Opus 5 tops an ARC-AGI-3 chart but at much higher evaluation cost

Angaisb_ · x · 2026-07-25

The attached chart compares Opus 5, Opus 4.8, and GPT-5.6 Sol on ARC-AGI-3 “novel problem-solving by cost.” Opus 5 sits far above the other points on the score axis, but at a much higher total evaluation cost than earlier runs. The image is essentially a benchmark snapshot arguing that Opus 5’s reasoning gains come with steep eval cost.

Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(33 posts)→

Original post →

More from Models

Models channel →