DeepSeek V4.1 'causal encoder-decoder' label debated: likely a YOCO-style KV-reuse design
donglixp · x · 2026-09-10
Developer aHapBean pushes back on the 'Causal Encoder-Decoder' framing of DeepSeek V4.1, arguing it overstates the architectural departure:
- V4.1 is likely still a decoder-only LLM; a clearer description is a YOCO-style KV-reuse design (You Only Cache Once)
- The 'encoder' is just the lower half of a jointly trained causal Transformer stack — early layers of decoder-only LLMs already function as encoders building contextual representations
- Upper layers share global KV derived from midpoint representations, letting most prompt tokens skip upper-layer computation during prefill — a real efficiency win
- The whole stack is trained end to end; 'encoder-decoder' wrongly suggests two separate networks
Related event: Inside DeepSeek V4.1 Flash: YOCO at its core and KV cache reuse(7 posts)→
More from Models
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11
- Dev opinion: model differences shrink in a good harness; Grok 4.6 is good enough — gnukeith · 2026-09-11