Bug Hunt Bench: blind-graded leaderboard testing frontier coding models on 105 real planted bugs

PawelHuryn · x · 2026-10-08

The author launched Bug Hunt Bench, a persistent leaderboard for frontier coding models tested on 105 real planted bugs. Highlights: blind grading, one prompt per repo, ranked by bugs fixed; sortable by cost, time, turns, and coverage; supports multiple reasoning tiers and harnesses, with superseded historical runs excluded from the default view. Full run notes and data are public on GitHub (run-notes.md, runs.csv). Built by Pawel Huryn of The Product Compass newsletter.

Original post →

More from Models

Models channel →