DeepSeek V4.1-Flash: New Architecture, Beats GPT-5.6 Sol on DeepSWE at 94% Lower Output Price
kimmonismus · x · 2026-09-10
DeepSeek released V4.1-Flash just six weeks after July's V4-Flash, with a brand-new architecture:
- Architecture: Causal Encoder–Decoder with native visual understanding; 552B MoE params with 8B active during input and 16B during output. Cheaper large-input reads plus shared/compressed stored context cut memory — ideal for agents repeatedly reading code, docs, and tool results.
- Efficiency: KV-cache needs cut to ¼ (HBM) and ⅛ (SSD) of the previous generation.
- Benchmarks (DeepSeek's own tests): 74.2% on DeepSWE vs 54.4% for V4-Flash and 73.0% for GPT-5.6 Sol; matches or beats Sol on several coding/agent benchmarks.
- Pricing: -32% fresh input, -57% cached input, -9% output vs V4-Flash. Peak rate $1.20/M output tokens vs Sol's $20 (-94%), off-peak $0.60.
The author: DeepSeek is back in the price war — "this is the real moat," and weekly releases with major jumps are the new normal.
More from Infra
- Speculative decoding has evolved twice: from small-batch speedups to lifting MoE compute intensity — YouJiacheng · 2026-09-10
- Millie's ternary 35B MoE hits 56% on SWE-bench Verified, runs local coding agent on iPhone 17 Pro — MannyKayy · 2026-09-10
- Matt Barrie burned 4B tokens in a day, cut his bill 500-fold, and now worries about $5T in debt — gaganghotra_ · 2026-09-10
- Analyst: DeepSeek's latest change is a big win for token efficiency, moving toward OpenAI's regime — teortaxesTex · 2026-09-10
- Acellera tests 7 LLM+harness combos on drug discovery: one RTX 5090 holds up — gdefabritiis · 2026-09-10
- Mac mini tested: local 35B runtime hits Haiku-level scores but falls short for agents — PawelHuryn · 2026-09-10