Claude Sonnet 5.5 Lands #3 on Agent Arena at $2.74/Task, Just Off the Pareto Frontier

arena · x · 2026-10-03

Agent Arena — a leaderboard ranking 51 models on real-world agentic tasks across 2.15M sessions — shows Claude Sonnet 5.5 debuting at #3 with a median cost of $2.74/task and a +12.5% net improvement score. It delivers top-tier performance at a premium: sibling Claude Opus 5.5 scores higher (+13.82%) at just $1.58/task, keeping Sonnet 5.5 just off the Pareto frontier.

Current Pareto-optimal models include Claude Fable 5.1 (Max, +14.31% at $4.62), Claude Opus 5.5 (High), GPT 6.1 Sol (Max, +11.23% at $0.57), DeepSeek V4.1 Flash (MIT, +4.02% at $0.10) and Xiaomi's MiMo V2.6 Pro/Flash. The top 10 also features GPT 6 Astra, Gemini 4 Argon, Claude Opus 5 and Kimi K3.

Related event: GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard(5 posts)→

Original post →

More from Models

Models channel →