FULL STORY

Claude Sonnet 5.5: From Rumor to Benchmark

Rumors of Claude Sonnet 5.5 scoring 56 on the Artificial Analysis index preceded its release. Post-launch benchmarks confirmed near-Opus performance, with writing on par and record token usage at high effort.

2026-09-29 ~ 2026-09-29 · 3 episodes · 17 posts

Episode 1 · Rumor: Claude Sonnet 5.5 Scores 56 on Intelligence Index, Nearing Opus 5.5 (2026-09-29, 3 posts)

Leaks from netizens and LuminaBench claim Claude Sonnet 5.5 scores 56 on the Artificial Analysis Intelligence Index, ranking second behind Claude Opus 5.5 and ahead of GPT-6 Astra, with competitor Sol at 48. The figures remain unconfirmed.

Episode 2 · Sonnet 5.5 Nearly Matches Opus 5.5 but Sets Token Consumption Record (2026-09-29, 12 posts)

Following Anthropic's release of Claude Sonnet 5.5, third-party evaluator Artificial Analysis published first benchmark data: the model scored 56 on the AA Intelligence Index, just 2 points below Opus 5.5 (max), with particularly strong performance under max effort; its Terminal-Bench 4.0 score of 64% is a 50-point improvement over Sonnet 5 (max). But while approaching flagship performance, its token consumption and cost are notably high — the core controversy of this evaluation round.

Confirmed

  • Intelligence Index score of 56, only 2 points behind Opus 5.5 (max); AA has published complete breakdowns of Intelligence Index results across reasoning effort levels.
  • At max effort, it averages about 193,000 output tokens per task — the highest of all tested models and roughly 7x GPT-6 Astra; compared with 119,000 for Opus 5.5 max, far exceeding sibling models.
  • Terminal-Bench 4.0 score of 64%, up 50 percentage points from Sonnet 5 (max).
  • AA's model comparison page shows the family's five reasoning effort tiers (max/xhigh/high/medium and below) scoring 56/52/47/41 points, with per-task costs ranging from $0.41 to $7.60 — an 18x spread.

Why it matters

  • Sonnet 5.5's mid-tier positioning near flagship performance reflects Anthropic's tiered reasoning-effort strategy, but the heavy consumption means real-world costs could double: @haider1's cost analysis calls it a "token guzzler," with per-task costs roughly 2x GPT-6.
  • For token-billed developers, Intelligence Index and unit cost must be weighed together, and AA's tiered breakdown offers a direct reference for model selection.

Episode 3 · Hands-on: Sonnet 5.5 matches Opus in writing at half the token price (2026-09-29, 2 posts)

Anthropic's new Sonnet 5.5 delivers an Opus-level upgrade, matching or beating the flagship in writing tests by Dan Shipper and Every's team. It costs half the tokens, though high effort mode consumes tokens quickly.