Claude Opus 5.5 Burns 15.6M Tokens per Task, 37% More Than Opus 5

ArtificialAnlys · x · 2026-09-24

Artificial Analysis launched a Coding Agent leaderboard that combines three benchmarks with equal weight: DeepSWE v1.1 (113 software engineering tasks), Terminal-Bench 4.0 (66 agentic terminal tasks) and SWE-Atlas-QnA (124 technical Q&A tasks), scoring pass@1 averaged over three attempts and reporting wall time and API cost per task.

It also reports that Claude Opus 5.5 uses 15.6M tokens per task versus 11.4M for Opus 5 (+37%), with cached input rising from 10.9M to 14.6M and output more than doubling from 137k to 333k.

Related event: Opus 5.5 burns 15.6M tokens per task, costs $13.04 despite price cut(3 posts)→

Original post →

More from coding & agent

coding & agent channel →