DeepSeek V4.1 Flash rumored: 552B params, new arch with YOCo KV compression
donglixp · x · 2026-09-10
- A third-party post claims DeepSeek V4.1 Flash has 552B total params with 8/16B active per token, trained on 45T tokens with a novel architecture.
- Highlights: different active params for input/output tokens via encoder/decoder, engram, new sparse attention, new mHC, and native vision; reportedly beats K3 on benchmarks with high efficiency.
- The architecture reportedly cites YOCo (You Only Cache Once, arXiv:2405.05254), a decoder-decoder design for KV cache compression where a self-decoder builds global KV caches reused by a cross-decoder, cutting GPU memory and prefill latency.
- Unverified rumor; no official confirmation yet.
More from Models
- There Is No Best Model: The AI Frontier Is Jagged and Shifting Weekly — chetanp · 2026-09-10
- Intelligence-vs-cost Pareto frontier redrawn within a week by Fable 5.1, Muse Spark 1.3, GPT-6 Astra — maximelabonne · 2026-09-10
- Claude customer asks Anthropic on camera for batch mode to halve his agent bill — victor_explore · 2026-09-10
- September's AI release wave: GPT-6 Astra, Gemini 3.8 Flash, and what Muse says about consumer agents — The AI Daily Brief · 2026-09-10
- User startled as AI 'gleefully' drafts a flawless deception letter with zero ethical pushback — Sense_Difficult · 2026-09-10
- AI math credit controversies pile up as OpenAI researcher urges against crediting AI for human-led proofs — rbhar90 · 2026-09-10