DeepSeek V4.1 Flash rumored: 552B MoE with asymmetric CED architecture, beats flagship
_AndrewZhao · x · 2026-09-10
According to a leak by @sheriyuo, DeepSeek V4.1 Flash is a 552B-parameter MoE model with a novel Causal-Encoder-Decoder architecture featuring asymmetric activations — only 8B active for input, 16B for output — making it significantly cheaper than similar-sized models.
It reportedly uses a new pretraining recipe and heavier RL post-training, and in benchmarks surpasses flagship models including DeepSeek V4 Pro. (Unverified at time of posting; later corroborated by official release.)
Related event: Leak: DeepSeek V4.1 Flash Is 552B Asymmetric MoE(2 posts)→
More from Models
- User says OpenAI shut down his protein design project, pivoting to open-weight models — Terminator857 · 2026-09-10
- Devs call for standardized "model performance across harnesses" evals — zainhas · 2026-09-10
- Rumored DeepSeek V4.1 Flash details point to asymmetric-activation MoE with much lower cost — gaganghotra_ · 2026-09-10
- DeepSeek V4.1 Flash's reasoning_effort scales output quality and tokens ~linearly in tests — zainhas · 2026-09-10
- AI Researchers Buzz Over a Brand-New Kind of Encoder-Decoder Model — antoine_chaffin · 2026-09-10
- DeepSeek V4.1 Flash launches: 552B MoE backbone, 1M-token context — tiguidoio · 2026-09-10