Third-party 10-dim eval shows DeepSeek V4.1 big reliability gains over V4-Pro-0813
teortaxesTex · x · 2026-09-13
Blogger @servasyyai ran DeepSeek V4.1 Flash through a custom ten-dimension eval, comparing it with V4-Pro-0813. V4.1 shows markedly improved reliability, posting a modestly higher second-round pass rate, while V4-Pro collapsed due to scoring zero on coding. DeepSeek had claimed at launch that V4.1 Flash beats Claude Fable 5.1 and trails only GPT 6.0 Astra; this third-party run offers an independent check on those claims.
More from Models
- Reddit user calls out OpenAI models for botching basic percentage math — nutinahut · 2026-09-13
- "Open-source AI must win": OpenMed manifesto argues powerful intelligence can't stay in few hands — MaziyarPanahi · 2026-09-13
- Mathematician cracks Navier–Stokes with Codex, then OpenAI agents did it in 88 hours — rubenhassid · 2026-09-13
- Grok Bots coming to XChat: leak suggests bot DMs on X — nima_owji · 2026-09-13
- Nex-N2.5 Pro plays Pokémon across hundreds of steps, testing real computer-use endurance — alifcoder · 2026-09-13
- Fine-tuned open-source models cost 95% less and beat frontier models, claims dev — ayushtweetshere · 2026-09-13