DeepSeek releases V4.1 Flash: 552B MoE beats V4 Pro, cuts prices, retires flagship
DeepSeek · wechat · 2026-09-10
DeepSeek officially launched V4.1 Flash, the smallest model in its new architecture family: a 552B MoE with a novel asymmetric Causal-Encoder-Decoder design (8B input / 16B output activations) and native vision understanding.
- Benchmarks beat DeepSeek V4 Pro and other flagships; V4 Pro will be retired, with requests to deepseek-v4-pro routed to V4.1 Flash from Sept 14.
- KV cache shrunk 4x vs the previous generation (1/4 HBM, 1/8 SSD needs), 437x vs the first model, sharply cutting agent-serving costs.
- New pricing takes effect Sept 10 with off-peak rates at half price; API name is deepseek-flash.
- Tencent WorkBuddy, CodeBuddy and OpenCode are launch partners; the model is MIT-licensed with a tech report, and the team is seeking large-scale deployers (2k GPUs + storage cluster).
Related event: DeepSeek Unveils Open-Source V4.1-Flash MoE Model(28 posts)→
More from Infra
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- After Nvidia's Hugging Face buyout, devs call for a neutral alternative — hargup13 · 2026-09-10
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested — No_Run8812 · 2026-09-10