Claude Sonnet 5.5 Lands #3 on Agent Arena at $2.74/Task, Just Off the Pareto Frontier
arena · x · 2026-10-03
Agent Arena — a leaderboard ranking 51 models on real-world agentic tasks across 2.15M sessions — shows Claude Sonnet 5.5 debuting at #3 with a median cost of $2.74/task and a +12.5% net improvement score. It delivers top-tier performance at a premium: sibling Claude Opus 5.5 scores higher (+13.82%) at just $1.58/task, keeping Sonnet 5.5 just off the Pareto frontier.
Current Pareto-optimal models include Claude Fable 5.1 (Max, +14.31% at $4.62), Claude Opus 5.5 (High), GPT 6.1 Sol (Max, +11.23% at $0.57), DeepSeek V4.1 Flash (MIT, +4.02% at $0.10) and Xiaomi's MiMo V2.6 Pro/Flash. The top 10 also features GPT 6 Astra, Gemini 4 Argon, Claude Opus 5 and Kimi K3.
Related event: GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard(5 posts)→
More from Models
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03
- Grok users hit usage limits with no upgrade path — and the account link is permanent — JOBhakdi · 2026-10-03
- Vercel brings Jev to its AI SDK for Python with an experimental evaluate() API — cramforce · 2026-10-03
- ChatGPT Pro user: OpenAI force-reset my weekly quota with 55% left — Garbia · 2026-10-03
- Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67% — flowersslop · 2026-10-03
- User: after forced switch to Gemini, my phone assistant barely works with 30s delays — loonydan42 · 2026-10-03