Claude Opus 5 tops an Artificial Analysis coding-agent benchmark at 67
Hesamation · x · 2026-07-25
A repost reacts to an Artificial Analysis chart comparing coding-agent performance across several models and says Anthropic “really cooked” with Opus 5. The attached image shows Claude Opus 5 at the top of the chart with a composite pass@1 score of 67, tied with Codex GPT-5.6 Sol and ahead of Claude Opus 5 and other models on the same benchmark set.
More from Models
- CogSci talk links linguistic efficiency to how LLMs learn meaning — ChrisGPotts · 2026-07-25
- Opus 5 reportedly scores 42/42 on IMO 2026 without tools — Afinetheorem · 2026-07-25
- Quadrillion says Anthropic’s Opus 5 is faster than Opus 4.8 on hard ML workloads — igarciacamargo · 2026-07-25
- Google is lagging behind open-weight models on most benchmarks — burny_tech · 2026-07-25
- Charts Show Opus 5 Peaks in Coding Performance with Medium Thinking — dejavucoder · 2026-07-25
- Opus 5, hidden-rule inference, and J-space point to a new agent stack — imjustnewatai · 2026-07-25