FULL STORY

Indie Bug Hunt Benchmark Crowns Conflicting Winners

Distrusting public benchmarks, developer Paweł Huryn built a Bug Hunt Benchmark on real repos; early runs crowned Sonnet 5.5, while later tests showed GPT-6.1 Sol fixing the hardest bugs.

2026-10-07 ~ 2026-10-07 · 2 episodes · 5 posts

Episode 1 · Independent Bug Hunt Benchmark Ranks Latest AI Models (2026-10-07, 3 posts)

Paweł Huryn tested the latest models with his own Bug Hunt Benchmark (2 real repos, 105 bugs missed by frontier models) and found Sonnet 5.5 tops the leaderboard, Opus 5.5 matches Fable 5.1 at two-thirds the cost, and GPT-6.1 Sol is both stronger and about 10x cheaper than its predecessor.

Episode 2 · Custom Bug Hunt benchmark flips results: GPT-6.1 Sol recovers (2026-10-07, 2 posts)

Paweł Huryn's self-built Bug Hunt Benchmark, which tests models on finding and fixing hard bugs in real repositories, produced different conclusions from other leaderboards: GPT-6.1 Sol rebounded while Muse models remained the cheapest workhorse.