Opus 5.5 sets Bug Hunt Bench record with 800+ self-check turns
PawelHuryn · x · 2026-09-29
Pawel Huryn reports record-breaking behavior on the Bug Hunt Bench: Opus 5.5 (max) finished with 138–163 model calls, while Sonnet 5.5 (max) made 628 calls, wrote its report at the 66-minute mark, and kept self-verifying for 111 minutes across 818 turns — behavior never seen before on the bench. Sonnet 5 (max) needed 207 turns. Scores will be reported after repo 2 is judged, repeated to n=3.
Related event: Opus 5.5 Sets New Record on Bug Hunt Bench(2 posts)→
More from coding & agent
- Pluto launches in beta: a personal agent with inbox, memory and a computer — Rasmic · 2026-09-29
- Anthropic adds official eval-building and hillclimbing workflow to Claude Code skill — ClaudeDevs · 2026-09-29
- Anthropic engineer: don't run Sonnet at max effort — use Opus instead — edwinarbus · 2026-09-29
- Hugging Face agent attack postmortem: allowlists gate where agents go, not what they do — kimmonismus · 2026-09-29
- PR Council MCP: Open-Source Multi-Agent PR Review Playground for Agentic Engineering — mostly_deterministic · 2026-09-29
- Moda ships Linear integration, says frustration detector beats Opus 5.5 at 1/30 the cost — KlausCodes · 2026-09-29