Bug Hunt Bench: Frontier Coding Models Graded Blind on 105 Planted Real-Repo Bugs
PawelHuryn · x · 2026-09-10
Pawel Huryn (The Product Compass) launched Bug Hunt Bench, a blind-graded board where frontier coding models fix planted bugs in real repos — 105 total, one prompt per repo, continuously updated.
Views include a leaderboard ranked by bugs fixed, score vs cost (costs span 200x on a log axis), score vs time, and coverage. A model can appear multiple times across reasoning tiers or harnesses, with superseded runs kept out of the default view, plus PNG export and custom run selection.
Related event: Bug Hunt Bench: DeepSeek V4.1 Flash Tops Price-Performance(8 posts)→
More from Models
- ValsAI launches RSI Index, first third-party benchmark measuring how close AI is to self-improvement — JenniferHli · 2026-09-11
- Devin's New Model Verdict: Not a Benchmaxxer, a 'Killer Execution Model' at $20/Month — brandon_galang · 2026-09-11
- Business Insider Asked ChatGPT, Gemini, Claude and Grok How AI Could End Humanity — coinfanking · 2026-09-11
- Claims resurface that Moonshot's Kimi distilled from Claude raw CoTs — xuanalogue · 2026-09-11
- User switches back to GPT-5.6 Sol: barely uses quota and feels faster — CtrlAltDwayne · 2026-09-11
- Dev opinion: model differences shrink in a good harness; Grok 4.6 is good enough — gnukeith · 2026-09-11