Unverified DeepSeek-V4.1-Flash report: 552B MoE slashing KV cache for million-token agent workloads
burkov · x · 2026-09-11
Andriy Burkov says a DeepSeek-4.1-Flash technical report is now readable (unverified — treat with caution until officially confirmed). Per the description: the model targets the rising cost of long-horizon agent workloads, where long inputs inflate prefill compute and large KV caches strain HBM and storage bandwidth. The claimed techniques combine a Causal Encoder-Decoder layout, cross-layer KV and index reuse in Compressed Sparse Attention 2, FP4 main KV storage, and bounded replay for sliding-window states, with a 552B-parameter backbone pretrained on 45 trillion multimodal tokens followed by standard SFT — aiming to slash KV cache far below prior levels while improving reasoning, agent, and visual performance.
More from Models
- Only Muse Spark 1.3 and Fable 5.1 sit on the coding Pareto frontier — jyangballin · 2026-09-11
- Muse Spark 1.3 shines on brand-new, unoptimizable CursorBench 4.0 — jyangballin · 2026-09-11
- User slams Anthropic for blocking benign queries on ancient texts and recursive AI — NickPassig · 2026-09-11
- Codex cybersecurity work needs the Daybreak model to avoid safety-guardrail blocks — HankYeomans · 2026-09-11
- DeepSeek's new open-source model reportedly crushes GLM and Kimi at 4-10x lower prices — anselm · 2026-09-11
- GPT-6 Astra rebuilds Cessna 337 landing gear from a YouTube video; Fable 5.1 falls short — FinanceYF5 · 2026-09-11