Author corrects Bug Hunt Bench numbers, contrast even starker
PawelHuryn · x · 2026-09-29
Pawel Huryn corrects his earlier Bug Hunt Bench report: the turn counts he initially gave were wrong, and the contrast between models is even higher than described. He will report scores after repo 2 is judged and repeat experiments to n=3.
Related event: Opus 5.5 Sets New Record on Bug Hunt Bench(2 posts)→
More from Models
- Dev after 2 days: Claude is excellent, Codex great for long-horizon tasks but poorly designed — cneuralnetwork · 2026-09-29
- Opus 5.5 tops Drone-Bench and cheats far less than prior Claude models — scaling01 · 2026-09-29
- User claims 'Opus 5.5' turned a post on agent harnesses into an explainer video in one shot — alex_verem · 2026-09-29
- Engineer proud as Sonnet 5.5 scores 61.6% on chartography benchmark — echen · 2026-09-29
- ProgramBench multi-agent eval: Opus 5.5 fastest with a 5-agent team, Sonnet 5.5 with subagents — jyangballin · 2026-09-29
- Arrow 2 Telos tops Design Arena's SVG generation benchmark — AWizardWhoCodes · 2026-09-29