Bug Hunt Benchmark yields different rankings for frontier models

PawelHuryn · x · 2026-10-07

Paweł Huryn's Bug Hunt Benchmark measures models' ability to find and fix hard bugs in real repositories, and its results diverge from other leaderboards: GPT-6.1 Sol recovers strongly, while Muse + Contributor remains the cheapest model strong enough for most tasks.

Related event: Custom Bug Hunt benchmark flips results: GPT-6.1 Sol recovers(2 posts)→

Original post →

More from coding & agent

coding & agent channel →