Bug Hunt Benchmark retest: GPT-6.1 Sol recovers, Muse still cheapest strong agent

PawelHuryn · x · 2026-10-07

PawelHuryn reports different results from his Bug Hunt Benchmark, which tests whether models can find and fix hard bugs in real repos. In this run, GPT-6.1 Sol manages to recover, while Muse Spark 1.3 + Contributor remains the cheapest model that is strong enough for most tasks. The post is part of his thread recalculating Artificial Analysis benchmark costs at subscription prices.

Related event: Custom Bug Hunt benchmark flips results: GPT-6.1 Sol recovers(2 posts)→

Original post →

More from coding & agent

coding & agent channel →