Bug Hunt Bench tests frontier coding models on 105 real bugs, costs vary 200x

PawelHuryn · x · 2026-09-22

Pawel Huryn launched Bug Hunt Bench, grading frontier coding models on 105 real (non-planted) bugs from Xiaomi repos with blind cross-family judges against a secret answer key. Results (score/cost): Muse Spark 1.3 max 32.2/$18.11; GPT-5.6 Luna max 31.5/$2.82; Grok 4.7 xhigh 28.8/$22.89; MiMo-V2.6-Pro 22.7/$0.86; DeepSeek V4.1 Flash 21.7/$0.78; Gemini 3.8 Flash 18/$11.03. Costs span 200x, forming a strong Pareto frontier. Full notes and data are on the benchmark site and GitHub.

Related event: Bug Hunt Bench Tests Models on 105 Real Bugs, Cost Gap Nears 200x(3 posts)→

Original post →

More from coding & agent

coding & agent channel →