DeepSeek V4.1 Flash halves KV cache with YOCO-style decoder-decoder design
nrehiew_ · x · 2026-09-11
nrehiew explains the core design of DeepSeek V4.1 Flash: 20-layer encoder plus 20-layer decoder. Since prefill is expensive, the bottom half builds the KV cache, which is shared with upper half layers via per-layer projection — roughly a state read individually by later layers. Inspired by YOCO, hence "decoder-decoder" with both modules causally masked; this doesn't apply to the SWA layers.
Related event: DeepSeek V4.1 Flash Deep Dive: KV Cache Compression at the Frontier(9 posts)→
More from Infra
- Google Commits $15B to AI Infrastructure Buildout in Finland — LinkedInNews · 2026-09-11
- DeepSeek-V4.1-Flash hits Ollama: 552B MoE backbone with 1M context via KV cache compression — ollama · 2026-09-11
- DIY-friendly KiCad footprints for AI MELF resistors, milled at home — debreuil · 2026-09-11
- Inference providers barely break even: $10K revenue yields just $200 profit — metalvendetta · 2026-09-11
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11