Sonnet 5.5 effort levels tested: max scores 55.5 at $134.79 vs 9 for Sonnet 5
PawelHuryn · x · 2026-09-29
Pawel Huryn benchmarked all Sonnet 5.5 effort levels on his real-repo bug-fixing benchmark (n=2 averages):
- max: 55.5, 1,330 turns, 234.7 min, $134.79
- xhigh: 39, 588 turns, $60.30
- high: 32, 294 turns, $16.73
- medium: 18, 112 turns, $8.39
- low: 20, 119 turns, $5.75
Sonnet 5 (max) scored just 9 at $17.96. Overall: an incredibly motivated model — fast and cheap per turn, but it runs far more turns when needed, with max costing 25x+ the low setting.
Methodology: real bugs re-planted in real repos that frontier models missed in early 2026, blind calibrated judges from a different family, only truly fixed bugs count, answer key kept private.
Related event: Bug Hunt Bench Launches: Sonnet 5.5 Leads Frontier Coding Models(4 posts)→
More from coding & agent
- Sonnet 5.5 one-shots a full $100K/month app in a single prompt — PrajwalTomar_ · 2026-09-29
- Open-source project trains an LLM from scratch in PyTorch, 13M params on free Colab — thisguyknowsai · 2026-09-29
- Spent 3 hours debugging an API that died in 2021: when Cursor's hallucinated legacy code wins the boring way — Full_Collar9026 · 2026-09-29
- ComfyUI creator shares Claude+Qwen+Blender video workflow, still fights consistency — Laxus534 · 2026-09-29
- Why this team stopped letting LLMs do compliance math and added episodic memory — Brilliant_Tension_53 · 2026-09-29
- Building a code review agent that remembers what it has seen — manivarsha · 2026-09-29