105 real bugs benchmarked: Sonnet 5.5 max scores 55.5, beating GPT-6 Astra at 45
PawelHuryn · x · 2026-09-29
PawelHuryn ran a bug-hunting benchmark on 2 real repos with 105 planted bugs. Results: Sonnet 5.5 (max) 55.5, GPT-6 Astra (max) 45, GPT-5.6 Sol (max) 43.5, Fable 5.1 (max) 43, Opus 5.5 (max) 41.7, Sonnet 5.5 (xhigh) 39. The key finding: Sonnet 5.5 max wins by being the least lazy model — it used 1,330 turns (6× Astra's 222, 3× Opus 5.5's) to fix 10–14 more bugs. The xhigh mode halved turns but dropped below Opus 5.5. Runs were validated with mutation testing (400 mutants, 2h23m per repo); the author plans at least 2 repeats per config.
Related event: Bug Hunt Bench Tests Frontier Models on 105 Real Bugs(2 posts)→
More from coding & agent
- Third-party tests back Fo agent's claim of 2x task completion with 94% trust rate — gaganghotra_ · 2026-09-29
- Qwen Open-Sources QwenGyre RL Framework for xLong-Horizon Agent Training — Qwen · 2026-09-29
- Meme: Coercing Your AI Agent to Follow Your Terrible Plan — mike64_t · 2026-09-29
- Agent Memory Should Have an Expiration Date: A Six-Field Metadata Framework — Hairy-Difficulty-411 · 2026-09-29
- Dev Builds Remote MCP Bridge Leting ChatGPT Web Chat Control Your Local PC — ChoasMaster777 · 2026-09-29
- 99-second demo: controlled terminal + browser agent execution via MCP with approval boundaries — ImaginaryMachine9110 · 2026-09-29