Leaked benchmarks reveal DeepSeek V4.1 and V4.1 Flash with novel 552B architecture
On September 10, several bloggers collectively leaked what are claimed to be official benchmark results and technical details for DeepSeek V4.1 and V4.1 Flash; none of the information has been officially confirmed.
Confirmed
- All information is unverified rumor, mainly from bloggers such as teortaxesTex, eliebakouch, and ChrisGPT, whose posts corroborate each other.
- The model has 552B total parameters with a brand-new Causal-Encoder-Decoder architecture, 8B input activation and 16B output activation.
- Leaked scores: TerminalBench 4.0 at 31.2, CyberGym 88.1 (vs. 84.5 for GPT-5.6 Sol), HLE 63.9; ChrisGPT claims it beats GPT-5.6 Sol on multiple agentic/coding benchmarks.
- The technical report cited by eliebakouch shows V4.1 Flash is a natively vision-capable model trained on 45T tokens, with an architecture praised as the most novel in years.
Architecture and Engineering Highlights
- KV cache compressed to 890 bytes/token, under 1GB for a million tokens; compared with V4-Flash, HBM requirements are down another 3.9x and SSD requirements down 8x, still unsurpassed half a year after release (per teortaxesTex).
- According to teortaxesTex, V4.1 removes V4's stopgap patches: dropping HCA, generalizing CSA into a more universal primitive, eliminating the warmup stage of sparse attention, and adopting a simpler single-pass mHC.
Not Yet Confirmed
- All scores, specs, and architecture details lack official sources; the final release may differ.
Why It Matters
- If true, V4.1 directly challenges GPT-5.6 Sol on agentic/coding capability while slashing deployment memory requirements, with significant implications for inference costs in an open-source ecosystem.
2026-09-10 ~ 2026-09-10 · 12 related posts
- Episode 1: DeepSeek V4.1 Flash surfaces in beta with new multimodal architecture(2026-09-08, 18 posts)
- Episode 2: DeepSeek V4.1 Flash Tested: Blazing 350 Tokens/s but Still Experimental(2026-09-08, 2 posts)
- Episode 3: DeepSeek cuts V4-Flash API pricing, off-peak cache hits drop to 0.02 yuan(2026-09-08, 6 posts)
- Episode 4: DeepSeek V4.1 Flash beta shows big gains in vision and cybersecurity(2026-09-09, 3 posts)
- Episode 5: DeepSeek launches V4.1-Flash: 552B MoE with native vision, 1M context, MIT licensed(2026-09-09, 26 posts)
- Episode 6: DeepSeek V4.1 Flash Matches 98% of GPT-6 Astra at 1.4% Cost(2026-09-09, 2 posts)
- Episode 7: Leaked benchmarks reveal DeepSeek V4.1 and V4.1 Flash with novel 552B architecture(2026-09-10, 12 posts)
Primary sources
- Leaked DeepSeek V4.1 benchmarks show 552B new-architecture model hitting 63.9 HLE with tools — teortaxesTex ·
- DeepSeek V4.1 Flash hits 74.2 on DeepSWE, beating Opus 5 and Gemini 3.8 Flash; asymmetric MoE revealed — nrehiew_ ·
- DeepSeek V4.1 Flash leak: 552B MoE with 8/16B active, native vision, praised as most novel arch in years — eliebakouch ·
- [source] Leaked DeepSeek V4.1 benchmarks show 552B new-architecture model hitting 63.9 HLE with tools — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Leak: 552B New Architecture, HBM Needs Cut 3.9x, Huge Benchmark Gains — teortaxesTex · 2026-09-10
- DeepSeek slashes KV cache to 890 bytes/token, hinting at million-agent swarms — teortaxesTex · 2026-09-10
- DeepSeek V4.1 Flash rumored: 552B MoE with asymmetric CED architecture, beats flagship — _AndrewZhao · 2026-09-10
- DeepSeek V4.1 Flash benchmarks leak: beats GPT-5.6 Sol on agentic/coding tasks at $0.14/M input — ChrisGPT · 2026-09-10
- [source] DeepSeek V4.1 Flash leak: 552B MoE with 8/16B active, native vision, praised as most novel arch in years — eliebakouch · 2026-09-10
- Leak: DeepSeek V4.1 reportedly drops V4's HCA, generalizes CSA, removes sparse-attention warmup — zephyr_z9 · 2026-09-10
- [source] DeepSeek V4.1 Flash hits 74.2 on DeepSWE, beating Opus 5 and Gemini 3.8 Flash; asymmetric MoE revealed — nrehiew_ · 2026-09-10
- DeepSeek V4.1-Flash leak: 552B params claiming GPT-5.6-level scores via CED architecture — iScienceLuvr · 2026-09-10
- Rumored DeepSeek V4.1 Flash details point to asymmetric-activation MoE with much lower cost — gaganghotra_ · 2026-09-10
- 552B params but only 8B active: DeepSeek V4.1 Flash's wild efficiency numbers — Scobleizer · 2026-09-10
- DeepSeek V4.1-Flash: KV cache 437x smaller, ~40x cheaper than Claude Opus 4.8 — AdinaYakup · 2026-09-10