Third-party 10-dim eval shows DeepSeek V4.1 big reliability gains over V4-Pro-0813

teortaxesTex · x · 2026-09-13

Blogger @servasyyai ran DeepSeek V4.1 Flash through a custom ten-dimension eval, comparing it with V4-Pro-0813. V4.1 shows markedly improved reliability, posting a modestly higher second-round pass rate, while V4-Pro collapsed due to scoring zero on coding. DeepSeek had claimed at launch that V4.1 Flash beats Claude Fable 5.1 and trails only GPT 6.0 Astra; this third-party run offers an independent check on those claims.

Original post →

More from Models

Models channel →