Bug Hunt Bench tests frontier coding models on 105 real bugs, costs vary 200x
PawelHuryn · x · 2026-09-22
Pawel Huryn launched Bug Hunt Bench, grading frontier coding models on 105 real (non-planted) bugs from Xiaomi repos with blind cross-family judges against a secret answer key. Results (score/cost): Muse Spark 1.3 max 32.2/$18.11; GPT-5.6 Luna max 31.5/$2.82; Grok 4.7 xhigh 28.8/$22.89; MiMo-V2.6-Pro 22.7/$0.86; DeepSeek V4.1 Flash 21.7/$0.78; Gemini 3.8 Flash 18/$11.03. Costs span 200x, forming a strong Pareto frontier. Full notes and data are on the benchmark site and GitHub.
Related event: Bug Hunt Bench Tests Models on 105 Real Bugs, Cost Gap Nears 200x(3 posts)→
More from coding & agent
- Scheduled Agent Produces a Full Short Episode Weekly From One Story Beat — socialwithaayan · 2026-09-22
- One Line In, Finished Episode Out: An Agent Replaces a Three-Tool Video Workflow — socialwithaayan · 2026-09-22
- One-line input to finished episode: an agent pipeline built on CREAO — socialwithaayan · 2026-09-22
- One prompt, one weekly AI episode: automating a mini-movie pipeline with CREAO and Seedance 2.5 — kamathsblog · 2026-09-22
- Jev, from a ChatGPT researcher, claims 200x faster and 400x cheaper decisions than LLMs — bibryam · 2026-09-22
- Storewake MCP server puts App Store rankings, RevenueCat revenue and Apple Ads into your AI assistant for $19/mo — Pfernan95 · 2026-09-22