Is DeepSeek falling behind? v4 Pro trails Qwen and GLM in benchmarks
power97992 · reddit · 2026-09-03
A Reddit post questions DeepSeek's standing: after topping open leaderboards with v3 and v3.2, DeepSeek v4 Pro 0813 now underperforms Qwen 3.8 Next, GLM 5.3 Flash and other frontier open models in benchmarks, while v4 Flash is decent. The poster speculates v4 Pro wasn't trained to full potential, that KV cache compaction and hybrid attention may hurt performance, notes talent reportedly leaving for Xiaomi and other labs, and asks whether DeepSeek can catch up within a month or two.
More from Models
- AxiomProver tops LeanEval, the last unsaturated math formalization benchmark — BenBlaiszik · 2026-09-03
- Muse Spark 1.3 calls user 'Judah' then denies it, users report odd behavior — fragment_me · 2026-09-03
- Google AI Mode shows zero citations on high-level TOFU queries, SEO tests find — gaganghotra_ · 2026-09-03
- Fable 5.1 halves agent failure rate to 7% with 0.7% hallucinations, at 1.8x the cost — ryanshrout · 2026-09-03
- Seroter Daily #859: Gemini 3.8 Flash, agent telemetry, and 7 agent skill patterns — rseroter · 2026-09-03
- Insider leak: OpenAI's Astra tested as 'ultima-alpha' and 'vega-alpha' checkpoints — Ok_Display_3159 · 2026-09-03