DeepSeek-V4.1-Flash cuts KV cache 4x with split encoder-decoder architecture
bendee983 · x · 2026-09-17
- DeepSeek's V4.1-Flash is called "a masterpiece in AI engineering and research". The model is nearly twice as large as its predecessor, yet KV storage dropped from 3,514 bytes per token to 890, making long agent sessions far cheaper to hold.
- The headline feature is a split encoder-decoder architecture: prefill runs through half the layers while decode uses the full stack, separating reading from writing.
- Key numbers: 8B active params on input, 16B on output; cache hits at $0.006 per million tokens; benchmark score of 40 vs Gemini's 41; a million-token session fits under 1GB of global KV (predecessor needed 3.5GB); local window state is no longer persisted to disk and persistent SSD cache falls to about one-eighth.
- Verdict: expect V4.1-Flash, like prior DeepSeek releases, to set a new precedent for open AI architectures.
Related event: DeepSeek V4.1 Flash architecture reset cuts KV cache to a quarter(5 posts)→
More from Infra
- Huawei unveils Ascend 960 chips for 2027 and UnifiedBus linking one million processors — mark_k · 2026-09-17
- Leak: CXMT Supplies TSV Embedded DRAM Dies for Chinese cHBM, Stacking Done In-House or via Local OSAT — zephyr_z9 · 2026-09-17
- OpenAI's Astra gets 3x throughput on NVIDIA Vera Rubin, plus 2x more in 72 hours — MickeySteamboat · 2026-09-17
- Baseten adds server-side web search for open models, 15% lower latency — baseten · 2026-09-17
- Training video LoRAs on 2x RTX 5060 Ti: consumer multi-GPU feasibility — Inner_Employment_332 · 2026-09-17
- 15 Small Models That Beat Models 100x Their Size at One Task — bigaiguy · 2026-09-17