Claude Sonnet 5.5 debuts at #3 in Agent Arena, costing 73% more per task than Opus 5.5
arena · x · 2026-10-03
Claude Sonnet 5.5 (Max) debuted at #3 in the Agent Arena with a +12.5% net improvement score — 8.1 points above Claude Sonnet 5 (High), which ranks #13 at +4.4%.
By category, Sonnet 5.5 took #1 in Chat (+15.6%), beating Fable 5.1 (+11.49%) and Opus 5.5 (+10.29%).
The performance comes at a premium: its median cost is $2.74 per task, about 73% higher than #2 Claude Opus 5.5 (High) at $1.58 — which also scores higher. That tradeoff keeps Sonnet 5.5 just off the Agent Arena Pareto frontier. Anthropic models now hold all three top spots.
Related event: GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard(5 posts)→
More from Models
- Veteran engineer: AI coding now beats humans on quality, not just speed — facontidavide · 2026-10-03
- Grok users hit usage limits with no upgrade path — and the account link is permanent — JOBhakdi · 2026-10-03
- Vercel brings Jev to its AI SDK for Python with an experimental evaluate() API — cramforce · 2026-10-03
- ChatGPT Pro user: OpenAI force-reset my weekly quota with 55% left — Garbia · 2026-10-03
- Dev vibecodes AtlasBench Europe spatial reasoning benchmark; GPT-6.1 tops at 84.67% — flowersslop · 2026-10-03
- User: after forced switch to Gemini, my phone assistant barely works with 30s delays — loonydan42 · 2026-10-03