Sonnet 5.5 effort levels tested: max scores 55.5 at $134.79 vs 9 for Sonnet 5

PawelHuryn · x · 2026-09-29

Pawel Huryn benchmarked all Sonnet 5.5 effort levels on his real-repo bug-fixing benchmark (n=2 averages):

Sonnet 5 (max) scored just 9 at $17.96. Overall: an incredibly motivated model — fast and cheap per turn, but it runs far more turns when needed, with max costing 25x+ the low setting.

Methodology: real bugs re-planted in real repos that frontier models missed in early 2026, blind calibrated judges from a different family, only truly fixed bugs count, answer key kept private.

Related event: Bug Hunt Bench Launches: Sonnet 5.5 Leads Frontier Coding Models(4 posts)→

Original post →

More from coding & agent

coding & agent channel →