Claude Opus 5 climbs to No. 2 in Agent Arena, ahead of GPT-5.6 Sol
scaling01 · x · 2026-07-29
Claude Opus 5 lands near the top of Agent Arena, ahead of GPT-5.6 Sol
A repost of Arena results says Anthropic’s Claude Opus 5 Max is now #2 in Agent Arena, with Opus 5 High at #3, based on more than 7,000 real-world agent sessions.
- Opus 5 Max: #2, net improvement 11.88%
- Opus 5 High: #3, net improvement 11.73%
- Both versions sit just behind Fable 5 and ahead of GPT-5.6 Sol
- Arena says the benchmark covers long-horizon tasks with web, filesystem, and terminal access
- The cited performance reflects outcome-based rankings from real agent workflows
Related event: Agent Arena: Claude Opus 5 Outperforms GPT-5.6 Variants(4 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23