105 real bugs benchmarked: Sonnet 5.5 max scores 55.5, beating GPT-6 Astra at 45

PawelHuryn · x · 2026-09-29

PawelHuryn ran a bug-hunting benchmark on 2 real repos with 105 planted bugs. Results: Sonnet 5.5 (max) 55.5, GPT-6 Astra (max) 45, GPT-5.6 Sol (max) 43.5, Fable 5.1 (max) 43, Opus 5.5 (max) 41.7, Sonnet 5.5 (xhigh) 39. The key finding: Sonnet 5.5 max wins by being the least lazy model — it used 1,330 turns (6× Astra's 222, 3× Opus 5.5's) to fix 10–14 more bugs. The xhigh mode halved turns but dropped below Opus 5.5. Runs were validated with mutation testing (400 mutants, 2h23m per repo); the author plans at least 2 repeats per config.

Related event: Bug Hunt Bench Tests Frontier Models on 105 Real Bugs(2 posts)→

Original post →

More from coding & agent

coding & agent channel →