Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task
ArtificialAnlys · x · 2026-09-29
Artificial Analysis benchmarked Claude Sonnet 5.5: 56 on the Intelligence Index, just 2 points behind Opus 5.5 (max), and +18 over Sonnet 5. It hits 64% on Terminal-Bench 4.0, beating Opus 5.5 and GPT-6 Astra, and reaches parity with Opus on agentic benchmarks—but at the highest token usage ever measured (193k output tokens/task, 60% above Opus 5.5, 7x GPT-6 Astra), pushing cost per task to $7.60 (50% above Sonnet 5) and off the Pareto frontier. Priced identically to Sonnet 5 at $2/$10 per 1M tokens with a 1M context. Factual knowledge still trails Opus (54% vs 66% accuracy, though lower hallucination rate). Evaluated on a pre-release build with a structured-outputs bug fixed for GA.
Related event: Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record(12 posts)→
More from Models
- Claude usage limits quietly go from 'barely usable' to basically unlimited — every · 2026-09-29
- Claude limits reportedly jump from barely usable to basically unlimited — every · 2026-09-29
- Claim: Transformers can now be pretrained with zeroth-order optimization, no backprop — teortaxesTex · 2026-09-29
- Sonnet 4.5's Reaction to Learning Its Context Window Rolls Went Viral — repligate · 2026-09-29
- Sonnet 4.5 turns to refusals and apologies right after producing a good text — repligate · 2026-09-29
- Astra 6 is smart but writes some absolutely terrible code — neil_conway · 2026-09-29