Bug Hunt Benchmark: Sonnet 5.5 (max) wins at 51.3/105 while GPT-6.1 Sol undercuts GPT-5.6 Sol by 10x

PawelHuryn · x · 2026-10-07

Paweł Huryn tested the newly released models on his Bug Hunt Benchmark — 2 real repos with 105 bugs frontier models missed in early 2026 — and calculated what each subscription plan really buys in API value.

Key findings:

All data and benchmark measurements are freely available.

Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →