DeepSeek launches V4.1-Flash: 552B MoE with native vision, 1M context, MIT licensed
On September 10, DeepSeek officially released and open-sourced its new model V4.1 Flash, rolling out simultaneously on web, app, and API, with the API model named deepseek-flash. This is one of the most notable new open-source model releases right now: performance reportedly surpasses their own flagship V4 Pro across the board, at a sharply lower price.
Confirmed
- Architecture and scale: a 552B total-parameter asymmetric MoE model with a brand-new Causal-Encoder-Decoder architecture, activating only 8B parameters on the input side and 16B on the output side; supports up to 1 million tokens of context, with KV cache overhead cut by 7/8.
- Benchmarks: per @ChrisGPT, V4.1 Flash beats GPT-5.6 Sol and Opus 5 on multiple agentic/coding benchmarks, e.g., CyberGym 88.1 vs 84.5; input pricing as low as $0.14/M.
- V4 Pro retirement and routing: the official announcement states that since V4.1 Flash surpasses V4 Pro across performance, cost, speed, and total time, all V4 Pro requests will be automatically routed to V4.1 Flash and billed at the lower price until V4.1 Pro launches. Several users (e.g., @gaganghotra) have already received deprecation notices.
- Older models sunset: the official API announcement discloses that legacy models like V4-Flash and V4-Flash-Vision-E will be retired.
Why it matters
- V4.1 Flash matches and exceeds flagship V4 Pro performance with far fewer activated parameters and much lower cost—combined with its open-source release, this could further drive down industry inference prices and pressure closed-model vendors.
- V4 Pro users should watch for behavior changes from the automatic routing; developers relying on its capabilities should evaluate V4.1 Flash's actual performance before switching.
- The 1M-token context and 7/8 KV cache savings have direct implications for usability and cost in long-document and agentic workflow scenarios.
2026-09-09 ~ 2026-09-10 · 26 related posts
- Episode 1: DeepSeek V4.1 Flash surfaces in beta with new multimodal architecture(2026-09-08, 18 posts)
- Episode 2: DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental(2026-09-08, 2 posts)
- Episode 3: DeepSeek cuts V4-Flash API pricing, off-peak cache hits drop to 0.02 yuan(2026-09-08, 6 posts)
- Episode 4: DeepSeek V4.1 Flash beta shows big gains in vision and cybersecurity(2026-09-09, 3 posts)
- Episode 5: DeepSeek launches V4.1-Flash: 552B MoE with native vision, 1M context, MIT licensed(2026-09-09, 26 posts)
- Episode 6: DeepSeek V4.1 Flash Matches 98% of GPT-6 Astra at 1.4% Cost(2026-09-09, 2 posts)
- Episode 7: Leaked benchmarks reveal DeepSeek V4.1 and V4.1 Flash with novel 552B architecture(2026-09-10, 12 posts)
Primary sources
- [source] DeepSeek routes V4 Pro requests to faster, cheaper V4.1 Flash, hints V4.1 Pro is coming — zephyr_z9 · 2026-09-09
- DeepSeek releases V4.1 Flash: 22% cheaper than V4 and better performance — deliprao · 2026-09-09
- DeepSeek ships V4.1 Flash: 552B MoE on new architecture, MIT-licensed — deepseek-ai · 2026-09-10
- DeepSeek releases V4.1 Flash: 552B MoE beats V4 Pro, cuts prices, retires flagship — DeepSeek · 2026-09-10
- DeepSeek Launches V4.1-Flash, Smallest Model in New Architecture Family with Native Vision — deepseek_ai · 2026-09-10
- DeepSeek Details V4.1-Flash: 552B MoE with Asymmetric 8B/16B Active Params, Beats Its Flagship — deepseek_ai · 2026-09-10
- [source] DeepSeek-V4.1-Flash Hits the API: 1/8 the SSD Cache, V4-Pro Being Phased Out — deepseek_ai · 2026-09-10
- DeepSeek Open-Sources V4.1-Flash Under MIT License, Cuts Off-Peak API Rates to 50% — deepseek_ai · 2026-09-10
- DeepSeek V4.1 Flash drops: 552B MoE activating just 8B, beating GPT-5.6 Sol at $0.14/M — ChrisGPT · 2026-09-10
- DeepSeek to retire v4 Pro, reroute all requests to new v4.1 Flash — gaganghotra_ · 2026-09-10
- DeepSeek releases V4.1 Flash: 552B MoE, 1M context, SGLang day-0 support — BanghuaZ · 2026-09-10
- Inside V4.1 Flash: Per-Token KV Cache Down to 890 Bytes, Peak-Valley Pricing — 赛博禅心 · 2026-09-10
- DeepSeek-V4.1-Flash appears on Hugging Face: MIT-licensed multimodal model with FP8 weights — AIFlow_ML · 2026-09-10
- DeepSeek V4.1 Flash weights go live on Hugging Face under MIT license — solyarisoftware · 2026-09-10
- DeepSeek V4.1 Flash launches: 552B MoE backbone, 1M-token context — tiguidoio · 2026-09-10
- DeepSeek V4.1 Flash now live on App, Web and API with open weights — gaganghotra_ · 2026-09-10
- Perplexity CEO reacts to DeepSeek V4.1 Flash launch — keviv9 · 2026-09-10
- DeepSeek V4.1 Flash debuts: 769B MoE with native vision, 1M context, day-0 vLLM support — solyarisoftware · 2026-09-10
- DeepSeek previews new architecture: 552B MoE with just 8B active input parameters — zephyr_z9 · 2026-09-10
7 near-duplicate retellings: 机器之心 · eliebakouch · reach_vb · teortaxesTex · basedjensen · Scobleizer · Scobleizer