Bug Hunt Bench updates all effort tiers; GPT-6.1 Sol xhigh nearly free on 105 planted bugs

PawelHuryn · x · 2026-09-30

Pawel Huryn's Bug Hunt Bench — blind-graded runs of frontier coding models fixing 105 real planted bugs across repos — has all effort levels ready, with n=3 additional runs in progress. xhigh-tier results are out, with GPT-6.1 Sol at xhigh virtually free. Full leaderboard, per-run notes and caveats live on bughunt.productcompass.pm and its GitHub (results/run-notes.md, runs.csv).

Related event: GPT-6.1 Sol benchmarks land: near-Astra performance at a fraction of the cost(16 posts)→

Original post →

More from Models

Models channel →