Bug Hunt Bench: Frontier Coding Models Graded Blind on 105 Planted Real-Repo Bugs

PawelHuryn · x · 2026-09-10

Pawel Huryn (The Product Compass) launched Bug Hunt Bench, a blind-graded board where frontier coding models fix planted bugs in real repos — 105 total, one prompt per repo, continuously updated.

Views include a leaderboard ranked by bugs fixed, score vs cost (costs span 200x on a log axis), score vs time, and coverage. A model can appear multiple times across reasoning tiers or harnesses, with superseded runs kept out of the default view, plus PNG export and custom run selection.

Related event: Bug Hunt Bench: DeepSeek V4.1 Flash Tops Price-Performance(8 posts)→

Original post →

More from Models

Models channel →