Bug Hunt Bench: New Blind-Graded Benchmark Tests Frontier Models on 105 Real Bugs

PawelHuryn · x · 2026-09-29

Pawel Huryn launched Bug Hunt Bench, a blind-graded benchmark testing frontier coding models on 105 real, re-planted bugs from real repos that frontier models missed in early 2026. Judges come from different model families, only fully fixed bugs score, and the answer key stays private to prevent gaming. The leaderboard also plots score against cost — runs span roughly 200x in cost — with full run notes on GitHub.

Related event: Bug Hunt Bench Launches: Sonnet 5.5 Leads Frontier Coding Models(4 posts)→

Original post →

More from Models

Models channel →