Opus 5.5 nearly matches Fable 5.1 at two-thirds the cost in real-repo bug benchmark

PawelHuryn · x · 2026-10-07

Distrusting public leaderboards, Paweł Huryn tested Opus 5.5, Sonnet 5.5 and GPT-6.1 Sol on his Bug Hunt Benchmark (2 real repos, 105 bugs frontier models missed in early 2026):

All data and measurements are freely available.

Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →