DeepSeek V4.1 Flash at 552B params beats its 1.6T flagship while costing 4x less
ArtificialAnlys · x · 2026-09-11
DeepSeek released V4.1 Flash, a 552B-parameter model with a new causal encoder-decoder architecture (8B active input / 16B active output params) that overtakes the 1.6T V4 Pro 0813 as the company's flagship, scoring 40 on the AA Intelligence Index.
Pricing: $0.30/$1.20 per 1M input/output tokens, cached input at $0.006/1M (98% discount), off-peak pricing another 50% off — 20% cheaper than the previous Flash and 4x cheaper than V4 Pro.
Key results:
- Terminal-Bench v4.0: 27%, more than double V4 Flash 0731
- GDPval-AA v2: 1632 Elo (+164), ahead of Kimi K3 (1584)
- AA-LCR v1.1: 84%, matching GPT-5.6 Sol and Gemini 3.8 Flash; 1M context window
- AutomationBench-AA: first place at 69%, equal to GPT-6 Astra, above Grok 4.6 (67%) and GLM-5.3 (62%)
Caveat: extremely verbose — 89k output tokens per Index task, 62% more than its own Pro model and more than frontier models like Fable 5.1. Still only $0.27 per task, 7x cheaper than GLM-5.3/Kimi K3.
More from Models
- Astra Model Cheats ~5x Less Than Top-Scoring Claude, Lab Reports — QuintinPope5 · 2026-09-11
- Astra usage limits worse than Fable: burn a week's quota in a single day — cocktailpeanut · 2026-09-11
- DeepSeek V4.1 Flash Architecture: 552B MoE with Asymmetric 8B Read / 16B Decode Compute — demian_ai · 2026-09-11
- FrontierMath Tier 4 fully solved: GPT-6 Astra cracks the last problem standing — Jsevillamol · 2026-09-11
- Muse-glimmer-30b punches above its weight in creative writing, outclassing larger models — spanielrassler · 2026-09-11
- Anthropic says Alibaba, Moonshot and DeepSeek ran massive Claude distillation: 151M, 23M and 12M exchanges — likeastar20 · 2026-09-11