Claude Opus 5 reaches 30.2% on ARC-AGI-3, far above prior frontier scores
EricBuess · x · 2026-07-25
A repost says ARC-AGI-3 was designed to be brutally hard: humans solve all environments, while as of March every frontier model was below 1%.
The attached chart shows Claude Opus 5 (high) reaching 30.2%, far above the previous reported frontier result of 7.8%. It also visualizes the cost-to-score tradeoff, with Opus 5 sitting near the high-cost, high-score end of the curve.
Related event: Claude Opus 5 Sets New SOTA on ARC-AGI-3 with Algebraic Reasoning(6 posts)→
More from Models
- Claude Opus 5 reportedly beats Fable 5 on a hard 3D coding test at 75% of the price — rohanpaul_ai · 2026-07-25
- Frontier lab rumor says Opus 5 ARC-AGI 3 score looks fake — flowersslop · 2026-07-25
- Chart says Claude Opus 5 blocks far less defensive coding than Fable 5 — repligate · 2026-07-25
- Claude Opus 5 lands, with DirectTerminal bringing richer Claude Code output to the terminal — draginol · 2026-07-25
- A Codex reset calendar shows usage limits do not always refresh at midnight UTC — petrusenko_max · 2026-07-25
- Grok 4.5 tops Ramp’s invoice test on 150,000 real business bills — elonmusk · 2026-07-25