Sonnet 5.5 wins by running 7x more turns than GPT-6 Astra, but costs most and is slowest

PawelHuryn · x · 2026-10-08

Follow-up in Pawel Huryn's coding benchmark thread: full score-vs-time, score-vs-cost, and score-vs-turns charts across all effort levels. Sonnet 5.5 (max) still leads, but only by running 7x more turns than GPT-6 Astra — the most persistent model he's benchmarked, yet also the most expensive and slowest on the board. He calls it likely a miscalibration by Anthropic that makes it impractical.

Related event: Pawel Huryn's Bug Hunt Benchmark Puts New Coding Models to the Test(5 posts)→

Original post →

More from Models

Models channel →