Claude Opus 5 reaches 30.2% on ARC-AGI-3, far above prior frontier scores

EricBuess · x · 2026-07-25

A repost says ARC-AGI-3 was designed to be brutally hard: humans solve all environments, while as of March every frontier model was below 1%.

The attached chart shows Claude Opus 5 (high) reaching 30.2%, far above the previous reported frontier result of 7.8%. It also visualizes the cost-to-score tradeoff, with Opus 5 sitting near the high-cost, high-score end of the curve.

Related event: Claude Opus 5 Sets New SOTA on ARC-AGI-3 with Algebraic Reasoning(6 posts)→

Original post →

More from Models

Models channel →