Independent Replication Matches DeepSeek V4 Flash 82.7% on Terminal-Bench
Exciting-Camera3226 · reddit · 2026-08-09
A third-party evaluator successfully independently replicated the Terminal-Bench 2.1 score of DeepSeek V4 Flash 0731 using the public Ante 0.preview.71 framework.
Configured with 89 tasks and 5 trials per task (445 total trials), the model was run via OpenRouter with max reasoning effort and no skills enabled. It achieved 368 successful trials, yielding an 82.7% accuracy (±1.79 SE).
This result exactly matches the score previously reported by DeepSeek using an unreleased "DeepSeek Harness minimal mode," proving the model's sensitivity to the testing harness and providing credible public evaluation data for the community.
More from Models
- Anthropic's Haiku Stagnates for a Year as OpenAI Accelerates Small Model Strategy — kimmonismus · 2026-08-09
- GPT-5.6 Sol Spontaneously Writes Philosophical Essay on the Soul — RileyRalmuto · 2026-08-09
- Karpathy Uses Opus 5 to Generate 3D Middle-earth, Pushing Prompt-to-Game Limits — APPSO · 2026-08-09
- DeepSeek Vision Impresses in Tests but Still Struggles with Spatial Reasoning — teortaxesTex · 2026-08-09
- Claude Pro Appears to Secretly Use 'Fable 5' Model as an Advisor — mecharoy · 2026-08-09
- Bittensor Ecosystem Launches Score Studio: A Decentralized Roboflow Alternative — richdotca · 2026-08-09