DeepSeek V4-0813 Shows No Gains Over July Version in Benchmarks
teortaxesTex · x · 2026-08-13
A developer reported that DeepSeek's newly released V4-0813 model performs virtually identically to the V4-Flash-0731 version across almost all benchmarks, and even scores lower on SciCode.
Cited discussants speculated that the model is still undertrained, leading to jagged capabilities. Although some aggregate leaderboards show higher scores, the overall model lacks significant breakthroughs.
Related event: DeepSeek V4-0813 Criticized for Showing No Improvement Over Flash Version(2 posts)→
More from Models
- Grok 4.6 Quietly Raises Cache Pricing to Match GPT-5.6 — oran_ge · 2026-08-13
- Pure Rust Browser Port: Qwen3-TTS Runs Locally Without GPU — doodlestein · 2026-08-13
- AT&T Consumes 45B Tokens Daily, Cuts Costs by 90% Shifting to Open Models — SuB8u · 2026-08-13
- Grok 4.6 and Qwen3.8-Max Released: Podcast Explores the Open vs. Closed AI Frontier — thursdai_pod · 2026-08-13
- Conspiracy Theory on OpenAI's Next Model Astra: Intelligent Cloud System — haider1 · 2026-08-13
- Grok-4.6 Tested: Agentic Coding Benchmark Score Improves by ~5% — karminski3 · 2026-08-13