Bug Hunt Benchmark: GPT-6.1 Sol matches or beats GPT-5.6 Sol at over 10x lower cost

PawelHuryn · x · 2026-10-07

Paweł Huryn tested Opus 5.5, Sonnet 5.5, and GPT-6.1 Sol on his Bug Hunt Benchmark: 105 bugs that frontier models missed in early 2026, across 2 real repos. Key findings:

All data and benchmark measurements are free to read.

Related event: Independent Bug Hunt Benchmark Ranks Latest AI Models(3 posts)→

Original post →

More from coding & agent

coding & agent channel →