Claude Opus 5.5 Tops Coding Agent Index but Costs 21% More Per Task
Artificial Analysis has launched its Coding Agent benchmark leaderboard, where Claude Opus 5.5 took the top spot with a score of 66 — the highest they've ever measured — even as token consumption and real cost per task both went up.
Confirmed
- Artificial Analysis launched its Coding Agent Index, with benchmarks including DeepSWE v1.1 (113 software engineering tasks), Terminal-Bench 4.0 (66 terminal agent tasks), and SWE-Atlas-QnA (124 tasks).
- Claude Opus 5.5 scored 66 in Claude Code's max-effort setting, ranking first and beating Opus 5's 60 — the highest score the firm has measured.
- Opus 5.5 consumed 15.6M tokens per task, 37% more than Opus 5's 11.4M; cached input rose from 10.9M to 14.6M, and output tokens roughly doubled from 137K to 330K.
- Even though Anthropic cut token prices at the same time, Opus 5.5's cost per task still came to $13.04, 21% more expensive than Opus 5.
Why it matters
- @ivanbezdomny relayed user feedback that the only clear downside of the new Opus 5.5 is extreme token hunger — its per-task token consumption on the Intelligence Index is the highest among all models tested; this may also explain Anthropic's pricing logic: cheaper per token, but heavier usage means the total bill isn't necessarily lower.
- For teams running coding agents at scale, "total cost per task" reflects real spending better than "price per token," and this leaderboard offers a direct reference for model selection.
2026-09-23 ~ 2026-09-24 · 5 related posts
- Episode 1: Claude Opus 5.5 Spotted in Claude Code Ahead of Imminent Launch(2026-09-23, 8 posts)
- Episode 2: Claude Opus 5.5 Launches Cheaper and Faster, Topping the AA Intelligence Index(2026-09-23, 195 posts)
- Episode 3: Perplexity Rolls Out Opus 5.5 to All Users, Citing 67.6% Cost Savings(2026-09-23, 2 posts)
- Episode 4: Anthropic Teases Sonnet 5.5 and Haiku 5.5 Within Weeks(2026-09-23, 2 posts)
- Episode 5: Claude Code 2.1.280 Ships with Opus 5.5 as Default Model(2026-09-23, 3 posts)
- Episode 6: Claude Opus 5.5 Rumored to Fall Back to Weaker Models on Frontier Dev Tasks(2026-09-23, 4 posts)
- Episode 7: Claude Opus 5.5 Tops Coding Agent Index but Costs 21% More Per Task(2026-09-23, 5 posts)
- Episode 8: Opus 5's FrontierCode Score Drops from 53.4% to 48% in New Card(2026-09-23, 3 posts)
- Episode 9: Opus 5.5 Wins Praise for Navier-Stokes Video Generation(2026-09-23, 2 posts)
- Episode 10: Developers Say Opus 5.5 Is Now Their Daily Driver: Faster and Cheaper(2026-09-23, 2 posts)
- Episode 11: Claude Opus 5.5 Wins Developer Praise, Overshadowing GPT-6 Sol(2026-09-23, 17 posts)
- Episode 12: Dev Rushes to Stress-Test Opus 5.5 Before Potential Nerf(2026-09-23, 2 posts)
Primary sources
- Claude Opus 5.5 tops the Coding Agent Index at 66, but cost per task jumps 21% — ArtificialAnlys ·
- Opus 5.5 uses 15.6M tokens per task, output more than doubles to 333k — ArtificialAnlys ·
- Cheaper tokens, pricier tasks: Opus 5.5 costs $13.04 per coding task, up 21% — ArtificialAnlys ·
- Opus 5.5 is a "token destroyer": highest per-task token usage of any benchmarked model — ivan_bezdomny · 2026-09-23
- [source] Claude Opus 5.5 tops the Coding Agent Index at 66, but cost per task jumps 21% — ArtificialAnlys · 2026-09-24
- [source] Cheaper tokens, pricier tasks: Opus 5.5 costs $13.04 per coding task, up 21% — ArtificialAnlys · 2026-09-24
- [source] Opus 5.5 uses 15.6M tokens per task, output more than doubles to 333k — ArtificialAnlys · 2026-09-24
- Claude Opus 5.5 Burns 15.6M Tokens per Task, 37% More Than Opus 5 — ArtificialAnlys · 2026-09-24