Artificial Analysis: Sonnet 5.5 Nears Flagship Performance but Burns the Most Tokens

Following Anthropic's release of Claude Sonnet 5.5, third-party evaluator Artificial Analysis published first benchmark data: the model scored 56 on the AA Intelligence Index, just 2 points below Opus 5.5 (max), with particularly strong performance under max effort; its Terminal-Bench 4.0 score of 64% is a 50-point improvement over Sonnet 5 (max). But while approaching flagship performance, its token consumption and cost are notably high — the core controversy of this evaluation round.

Confirmed

Why it matters

2026-09-29 ~ 2026-09-29 · 8 related posts

Full story(2 episodes)→

Primary sources