FULL STORY
Indie Bug Hunt Benchmark Crowns Conflicting Winners
Distrusting public benchmarks, developer Paweł Huryn built a Bug Hunt Benchmark on real repos; early runs crowned Sonnet 5.5, while later tests showed GPT-6.1 Sol fixing the hardest bugs.
2026-10-07 ~ 2026-10-07 · 2 episodes · 5 posts
Episode 1 · Independent Bug Hunt Benchmark Ranks Latest AI Models (2026-10-07, 3 posts)
Paweł Huryn tested the latest models with his own Bug Hunt Benchmark (2 real repos, 105 bugs missed by frontier models) and found Sonnet 5.5 tops the leaderboard, Opus 5.5 matches Fable 5.1 at two-thirds the cost, and GPT-6.1 Sol is both stronger and about 10x cheaper than its predecessor.
- Bug Hunt Benchmark: Sonnet 5.5 (max) wins at 51.3/105 while GPT-6.1 Sol undercuts GPT-5.6 Sol by 10x — PawelHuryn · 2026-10-07
- Opus 5.5 nearly matches Fable 5.1 at two-thirds the cost in real-repo bug benchmark — PawelHuryn · 2026-10-07
- Bug Hunt Benchmark: GPT-6.1 Sol matches or beats GPT-5.6 Sol at over 10x lower cost — PawelHuryn · 2026-10-07
Episode 2 · Custom Bug Hunt benchmark flips results: GPT-6.1 Sol recovers (2026-10-07, 2 posts)
Paweł Huryn's self-built Bug Hunt Benchmark, which tests models on finding and fixing hard bugs in real repositories, produced different conclusions from other leaderboards: GPT-6.1 Sol rebounded while Muse models remained the cheapest workhorse.
- Bug Hunt Benchmark yields different rankings for frontier models — PawelHuryn · 2026-10-07
- Bug Hunt Benchmark retest: GPT-6.1 Sol recovers, Muse still cheapest strong agent — PawelHuryn · 2026-10-07