Unverified DeepSeek-V4.1-Flash report: 552B MoE slashing KV cache for million-token agent workloads

burkov · x · 2026-09-11

Andriy Burkov says a DeepSeek-4.1-Flash technical report is now readable (unverified — treat with caution until officially confirmed). Per the description: the model targets the rising cost of long-horizon agent workloads, where long inputs inflate prefill compute and large KV caches strain HBM and storage bandwidth. The claimed techniques combine a Causal Encoder-Decoder layout, cross-layer KV and index reuse in Compressed Sparse Attention 2, FP4 main KV storage, and bounded replay for sliding-window states, with a 552B-parameter backbone pretrained on 45 trillion multimodal tokens followed by standard SFT — aiming to slash KV cache far below prior levels while improving reasoning, agent, and visual performance.

Related event: DeepSeek V4.1 Flash Architecture: A Flagship Model Built Around KV Cache Compression(10 posts)→

Original post →

More from Models

Models channel →