GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard

LMArena's leaderboard update this week brought significant shifts to the Agent Arena landscape: Anthropic's Claude Sonnet 5.5 (Max) debuted at #3 overall with a +12.5% net improvement score, while OpenAI's GPT-6.1 Sol (Max) landed at #5 with +11.23%, dramatically cutting the unit cost of frontier models — together reshaping the performance-cost Pareto frontier.

Confirmed

Why it matters

2026-10-03 ~ 2026-10-03 · 5 related posts

Primary sources