Sonnet 5.5 wins by running 7x more turns than GPT-6 Astra, but costs most and is slowest
PawelHuryn · x · 2026-10-08
Follow-up in Pawel Huryn's coding benchmark thread: full score-vs-time, score-vs-cost, and score-vs-turns charts across all effort levels. Sonnet 5.5 (max) still leads, but only by running 7x more turns than GPT-6 Astra — the most persistent model he's benchmarked, yet also the most expensive and slowest on the board. He calls it likely a miscalibration by Anthropic that makes it impractical.
Related event: Pawel Huryn's Bug Hunt Benchmark Puts New Coding Models to the Test(5 posts)→
More from Models
- Dev finds 6.1 Sol surprisingly good at designing native iOS apps — Dimillian · 2026-10-08
- Claude Code lead resets usage limits after community vote: users pick capacity over new features — CurieuxExplorer · 2026-10-08
- Prompting Opus 5.5 with a "grumpy senior engineer reviewer" kept it benchmarking for 2 days — remilouf · 2026-10-08
- Claude Haiku 5.5 beats GPT-6 Luna at matching price, but burns ~3x the tokens — Latent Space · 2026-10-08
- Gemini 3.8 Flash: 3 images and 5 chat messages already eat 8% of the quota — Horror-Airport-7606 · 2026-10-08
- Haiku 5.5 benchmarks 2-3x faster than GPT-6 Luna and performs better — PawelHuryn · 2026-10-08