Bug Hunt Bench: Sonnet 5.5 max scores 55.5 but costs $134.79 per run
PawelHuryn · x · 2026-09-29
Pawel Huryn added a log-scale score-vs-turns view to Bug Hunt Bench, his blind-graded benchmark of frontier coding models on 105 real planted bugs, and published per-effort numbers for Sonnet 5.5:
- max: 55.5, 1,330 turns, 234.7 min, $134.79 (n=2)
- xhigh: 39, 588 turns, $60.30
- high: 32, 294 turns, $16.73
- medium: 18, 112 turns, $8.39
- low: 20, 119 turns, $5.75
For comparison, Sonnet 5 max scored 9 at $17.96. The picture: an extremely motivated model, fast and cheap per turn, but cost scales steeply with effort — max is 3x the score of low at 25x the price. Cost spread across the board is roughly 200x.
Related event: Bug Hunt Bench Launches: Sonnet 5.5 Leads Frontier Coding Models(4 posts)→
More from coding & agent
- When 'Looks Correct' Becomes 'Verified': The LLM Coding Acceptance Problem — T_hompson · 2026-09-29
- A naive video transcription fixer pipeline: extract audio+frames, ASR, then correct with a frontier model — capetorch · 2026-09-29
- bdsqlsz is vibe-coding DLSS5 weight training for his in-development 3D game — bdsqlsz · 2026-09-29
- We Built an On-Call Agent That Failed the Right Way — Memory Can Learn the Wrong Lesson — Similar-Split7292 · 2026-09-29
- Extending Jev Mode to Images: Constrained llama.cpp Outputs as Image Selections — opUserZero · 2026-09-29
- Scraping Xiaohongshu hit posts with Codex + a wired Android phone — huangyun_122 · 2026-09-29