Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task

ArtificialAnlys · x · 2026-09-29

Artificial Analysis benchmarked Claude Sonnet 5.5: 56 on the Intelligence Index, just 2 points behind Opus 5.5 (max), and +18 over Sonnet 5. It hits 64% on Terminal-Bench 4.0, beating Opus 5.5 and GPT-6 Astra, and reaches parity with Opus on agentic benchmarks—but at the highest token usage ever measured (193k output tokens/task, 60% above Opus 5.5, 7x GPT-6 Astra), pushing cost per task to $7.60 (50% above Sonnet 5) and off the Pareto frontier. Priced identically to Sonnet 5 at $2/$10 per 1M tokens with a 1M context. Factual knowledge still trails Opus (54% vs 66% accuracy, though lower hallucination rate). Evaluated on a pre-release build with a structured-outputs bug fixed for GA.

Related event: Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record(12 posts)→

Original post →

More from Models

Models channel →