GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard
LMArena's leaderboard update this week brought significant shifts to the Agent Arena landscape: Anthropic's Claude Sonnet 5.5 (Max) debuted at #3 overall with a +12.5% net improvement score, while OpenAI's GPT-6.1 Sol (Max) landed at #5 with +11.23%, dramatically cutting the unit cost of frontier models — together reshaping the performance-cost Pareto frontier.
Confirmed
- Agent Arena is based on real Agent Mode tasks, with 2.15M+ cumulative sessions across 51 models; the live leaderboard and Pareto frontier data are published officially by Arena.
- Claude Sonnet 5.5 (Max) debuted at #3 overall with a net improvement of +12.5% and a median per-task cost of $2.74 — 8.1 percentage points higher than Claude Sonnet 5 (High) at #13 (+4.4%), but its per-task cost is 73% higher than Opus.
- GPT-6.1 Sol (Max) entered Agent Arena after OpenAI DevDay, ranking #5 with a +11.23% net improvement and a median per-task cost of $0.56, 39% lower than GPT-6 Sol (another post cited real-time data of $0.57/task).
- Models on the Pareto frontier also include Claude Fable 5.1 (Max), with a net improvement of +14.31% at $4.62/task.
- On this week's text leaderboard, Gemini 4 Argon took the top spot in Text Arena; another post noted Sonnet 5.5 closing in on GPT-6 at roughly a 20% price gap, with four new models collectively reshaping the Pareto frontiers of Text, Code, and Agent Arena.
Why it matters
- Sonnet 5.5 now performs at tier-one levels, but its $2.74 per-task cost makes its value proposition notably poor — a contrast with its price-competitive approach to GPT-6 in Code Arena, suggesting users must carefully weigh performance against cost.
- GPT-6.1 Sol entered the top five at under 40% of the cost, signaling that prices for frontier agent capabilities are falling fast, and the reshaped Pareto frontier could change how enterprises and developers choose models.
2026-10-03 ~ 2026-10-03 · 5 related posts
Primary sources
- Arena Weekly: Gemini 4 Argon Tops Text Arena, Sonnet 5.5 Within 2 Points of GPT-6 at 80% Less Cost — arena ·
- Agent Arena leaderboard: GPT-6.1 Sol joins Pareto frontier at $0.57/task, DeepSeek cheapest open model — arena ·
- Claude Sonnet 5.5 Lands #3 on Agent Arena at $2.74/Task, Just Off the Pareto Frontier — arena ·
- [source] Arena Weekly: Gemini 4 Argon Tops Text Arena, Sonnet 5.5 Within 2 Points of GPT-6 at 80% Less Cost — arena · 2026-10-03
- Claude Sonnet 5.5 debuts at #3 in Agent Arena, costing 73% more per task than Opus 5.5 — arena · 2026-10-03
- [source] Claude Sonnet 5.5 Lands #3 on Agent Arena at $2.74/Task, Just Off the Pareto Frontier — arena · 2026-10-03
- GPT-6.1 Sol hits #5 on Agent Arena, costs 39% less than GPT-6 Sol while scoring higher — arena · 2026-10-03
- [source] Agent Arena leaderboard: GPT-6.1 Sol joins Pareto frontier at $0.57/task, DeepSeek cheapest open model — arena · 2026-10-03