Opus 5.5 (High) hits #2 on Agent Arena, matching Opus 5 quality at 40-56% lower cost
arena · x · 2026-09-29
Agent Arena, a live agent-task leaderboard with 2M+ sessions, shows Claude Opus 5.5 (High) at #2, reshaping the Pareto frontier.
- Opus 5.5 (High): $1.31 median per task / +12.15% net improvement — beating both Opus 5 variants at 40-56% lower cost ($2.98 Max, $2.17 High).
- Other Pareto-optimal models: Claude Fable 5.1 (Max) +13.84% / $3.51, GPT 6 Sol (Max) +8.18% / $0.82, Tencent Hy4 preview +4.25% / $0.20, DeepSeek V4.1 Flash +3.96% / $0.07.
- The full ranking covers 45 models including Kimi K3, Grok 4.7 and Gemini 3.8 Flash, scored on tool reliability, task completion and steerability.
Related event: Claude Opus 5.5 Debuts Second on Agent Arena at 40-56% Lower Cost(2 posts)→
More from Models
- Tesla Roadster event pushed to Oct 15 over weather, with OpenAI set to announce tomorrow — Scobleizer · 2026-09-29
- Founder running 20+ startups says Opus 5.5 is AGI by his personal benchmarks — jonathan_wilke · 2026-09-29
- Opus 5.5 dramatically cuts em-dash usage, blurring AI writing detection — jonathan_wilke · 2026-09-29
- Sonnet 5.5 Sets Arena Record for Most Output Tokens, Sparking Pricing Debate — Gohab2001 · 2026-09-29
- PrivacyBench v2 launches: micro1's flow-transform 1.0 leads at 95.84%, 9.64 points ahead — Exp_Mark · 2026-09-29
- "Astra was incredible yesterday, terrible today": user reports overnight quality drop — HairyHobNob · 2026-09-29