Opus 5 tops an ARC-AGI-3 chart but at much higher evaluation cost

Angaisb_ · x · 2026-07-25

The attached chart compares Opus 5, Opus 4.8, and GPT-5.6 Sol on ARC-AGI-3 “novel problem-solving by cost.” Opus 5 sits far above the other points on the score axis, but at a much higher total evaluation cost than earlier runs. The image is essentially a benchmark snapshot arguing that Opus 5’s reasoning gains come with steep eval cost.

Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→

Original post →

More from Models

Models channel →