Custom Bug Hunt benchmark flips results: GPT-6.1 Sol recovers

Paweł Huryn's self-built Bug Hunt Benchmark, which tests models on finding and fixing hard bugs in real repositories, produced different conclusions from other leaderboards: GPT-6.1 Sol rebounded while Muse models remained the cheapest workhorse.

2026-10-07 ~ 2026-10-07 · 2 related posts

Full story(2 episodes)→