DeepSeek Details V4.1-Flash: 552B MoE with Asymmetric 8B/16B Active Params, Beats Its Flagship
deepseek_ai · x · 2026-09-10
DeepSeek shared the architecture behind V4.1-Flash, key to its claimed flagship-beating results:
- 552B-parameter MoE with a novel Causal Encoder–Decoder architecture, highly asymmetric activation: only 8B active parameters for input, 16B for output
- Combined with new pre-training methods and larger-scale RL post-training, benchmark results surpass flagship models including DeepSeek-V4-Pro
More from Models
- DeepSeek V4.1 Flash launches: 552B MoE backbone, 1M-token context — tiguidoio · 2026-09-10
- vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture — vllm_project · 2026-09-10
- DeepSeek V4.1 Flash ships under MIT license on HF and ModelScope — HugeConsideration211 · 2026-09-10
- DeepSeek V4.1 Flash benchmark chart circulates after release — toastisthicc · 2026-09-10
- Pro 5x user rage-quits over Codex usage limits, says Plus gets more coding done — DrunkenPionier · 2026-09-10
- DeepSeek V4.1 Flash rumored days away: 522B params, 1/4 KV cache, vision added — R_Duncan · 2026-09-10