Artificial Analysis: Sonnet 5.5 Nears Flagship Performance but Burns the Most Tokens
Following Anthropic's release of Claude Sonnet 5.5, third-party evaluator Artificial Analysis published first benchmark data: the model scored 56 on the AA Intelligence Index, just 2 points below Opus 5.5 (max), with particularly strong performance under max effort; its Terminal-Bench 4.0 score of 64% is a 50-point improvement over Sonnet 5 (max). But while approaching flagship performance, its token consumption and cost are notably high — the core controversy of this evaluation round.
Confirmed
- Intelligence Index score of 56, only 2 points behind Opus 5.5 (max); AA has published complete breakdowns of Intelligence Index results across reasoning effort levels.
- At max effort, it averages about 193,000 output tokens per task — the highest of all tested models and roughly 7x GPT-6 Astra; compared with 119,000 for Opus 5.5 max, far exceeding sibling models.
- Terminal-Bench 4.0 score of 64%, up 50 percentage points from Sonnet 5 (max).
- AA's model comparison page shows the family's five reasoning effort tiers (max/xhigh/high/medium and below) scoring 56/52/47/41 points, with per-task costs ranging from $0.41 to $7.60 — an 18x spread.
Why it matters
- Sonnet 5.5's mid-tier positioning near flagship performance reflects Anthropic's tiered reasoning-effort strategy, but the heavy consumption means real-world costs could double: @haider1's cost analysis calls it a "token guzzler," with per-task costs roughly 2x GPT-6.
- For token-billed developers, Intelligence Index and unit cost must be weighed together, and AA's tiered breakdown offers a direct reference for model selection.
2026-09-29 ~ 2026-09-29 · 8 related posts
- Episode 1: Rumor: Claude Sonnet 5.5 Scores 56 on Intelligence Index(2026-09-29, 3 posts)
- Episode 2: Artificial Analysis: Sonnet 5.5 Nears Flagship Performance but Burns the Most Tokens(2026-09-29, 8 posts)
Primary sources
- Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task — ArtificialAnlys ·
- Claude Sonnet 5.5 Burns ~193k Output Tokens Per Task, 7x More Than GPT-6 Astra — ArtificialAnlys ·
- Claude Sonnet 5.5 Efforts Span 18x Price Range: $0.41 to $7.60 Per Task on AA Index — ArtificialAnlys ·
- [source] Sonnet 5.5 scores 56 on AA Index, 2 points off Opus 5.5—but with record 193k tokens per task — ArtificialAnlys · 2026-09-29
- Sonnet 5.5 jumps 50 points to 64% on Terminal-Bench 4.0, topping Opus 5.5 — ArtificialAnlys · 2026-09-29
- [source] Claude Sonnet 5.5 Burns ~193k Output Tokens Per Task, 7x More Than GPT-6 Astra — ArtificialAnlys · 2026-09-29
- Artificial Analysis Publishes Full Eval Breakdown for Claude Sonnet 5.5 Across Reasoning Efforts — ArtificialAnlys · 2026-09-29
- [source] Claude Sonnet 5.5 Efforts Span 18x Price Range: $0.41 to $7.60 Per Task on AA Index — ArtificialAnlys · 2026-09-29
- Anthropic launches Claude Sonnet 5.5: 56 on AA Index, just 2 points behind Opus 5.5 — jarrodwatts · 2026-09-29
- Sonnet 5.5 is a token-hungry monster: 193k output tokens, 2x GPT-6's cost per task — haider1 · 2026-09-29
- AiBreakfast: Sonnet 5.5 now beats Opus 5.5 and Fable 5.1 overall — AiBreakfast · 2026-09-29