DeepSeek V4.1 Flash fixes 24/105 real bugs for $1.80 vs Opus 5's $51.33

PawelHuryn · x · 2026-09-10

Pawel Huryn tested frontier models on 105 hidden real bugs across 2 repos (find-and-fix), with results and costs:

Methodology FAQ: bugs are real, tough issues frontier models struggled with in early 2026; unplanted bugs don't count (OpenAI models report many real-but-irrelevant issues that would inflate scores); each model is judged by a different model family with confirmed calibration; answer keys aren't public, anonymized runs linked on GitHub. Luna (max) is called crazy good and cost-effective; Muse Spark 1.3 went untested due to failed payments in Europe and unreliable OpenRouter effort levels. Verdict: DeepSeek V4.1 Flash is a really strong everyday model at an extremely low cost.

Related event: DeepSeek V4.1 Flash Nears Frontier Model Performance at 1/100–1/30 the Cost in Independent Tests(8 posts)→

Original post →

More from Models

Models channel →