GPT-6 Sol hits #6 on Agent Arena with +7.7% net improvement at 56% less cost per task
arena · x · 2026-09-26
Agent Arena's new leaderboard puts GPT-6 Sol (Max) at #6 with +7.7% net improvement across 4,000+ real-world agentic sessions, up from #8 for its predecessor. Versus GPT 5.6 Sol (xHigh), it gains 1.5 points of net improvement at half the per-token price, and jumps to #4 on Confirmed Success (+11.4% vs +2.9%).
On the cost-performance Pareto frontier, GPT-6 Sol sits at $0.75 median cost/task — 56% cheaper than Claude Fable 5 (High) at $1.72 — while trailing it by only 0.6 percentage points. The arena measures tool orchestration on long-horizon tasks with web, filesystem, and terminal tools across 2M+ sessions.
More from coding & agent
- One agent writes the fix, another reviews it: a two-agent code review workflow in Slack — Al_Grigor · 2026-09-26
- Academic agent Memex upgraded to Opus 5.5: writing quality fixed, experience much better — arjunrajlab · 2026-09-26
- Dev uses open-source Ling-3.0-flash-VL to let AI redesign the foldable iPhone in a single HTML file — alifcoder · 2026-09-26
- Anthropic launches Claude plugin directory portal as MCP usage jumps 110x this year — ClaudeDevs · 2026-09-26
- Open-source Jev agent plays Pokemon Red live, pushing fast-decision AI beyond Tetris — supportingthedogs · 2026-09-26
- Anthropic deep dive: effort tuning in Claude Code pays off most for security and code review — trq212 · 2026-09-26