DeepSeek V4.1 Flash fixes 24/105 real bugs for $1.80 vs Opus 5's $51.33
PawelHuryn · x · 2026-09-10
Pawel Huryn tested frontier models on 105 hidden real bugs across 2 repos (find-and-fix), with results and costs:
- Opus 5 (max): 27 bugs, $51.33
- Grok 4.6 (max): 27 bugs, $16.96
- DeepSeek V4.1 Flash (max): 24 bugs, $1.80
- GPT-5.6 Luna (xhigh): 23 bugs, $2.50
- Opus 5 (high): 21 bugs, $38
Methodology FAQ: bugs are real, tough issues frontier models struggled with in early 2026; unplanted bugs don't count (OpenAI models report many real-but-irrelevant issues that would inflate scores); each model is judged by a different model family with confirmed calibration; answer keys aren't public, anonymized runs linked on GitHub. Luna (max) is called crazy good and cost-effective; Muse Spark 1.3 went untested due to failed payments in Europe and unreliable OpenRouter effort levels. Verdict: DeepSeek V4.1 Flash is a really strong everyday model at an extremely low cost.
More from Models
- DeepSeek's answer to surging demand: make its model cheaper and faster — yacineMTB · 2026-09-10
- Prelim assessment: GLM-4.1 behaves similar to V4, Zhipu's post-training called more advanced — menhguin · 2026-09-10
- Huge share of post-2022 web data is AI content mislabeled as human-written — menhguin · 2026-09-10
- DeepSeek V4.1 Hailed as the 'First Gamer Model' After Blowing Away an FPS-Generation Test — teortaxesTex · 2026-09-10
- Bug Hunt Bench grades frontier models on 105 real bugs; DeepSeek-V4.1-Flash lands 24/105 for $1.80 — PawelHuryn · 2026-09-10
- RSI is here, just disaggregated: DeepSeek using LLMs to design algorithms — teortaxesTex · 2026-09-10