Anthropic is using ARC-AGI-3 to gauge Opus 5’s novel problem-solving ability

typewriters · x · 2026-07-25

Anthropic is reportedly using ARC-AGI-3 to measure the novel problem-solving ability of Opus 5.

The attached chart compares models by total evaluation cost on the x-axis and score on the y-axis. In the visual, Opus 5 (high) is shown around 30% on ARC-AGI-3, while Opus 4.8 (high) sits near the low single digits and GPT-5.6 Sol is plotted below 10% at higher cost.

The post frames ARC-AGI-3 as a benchmark for frontier-style novel problem solving rather than routine benchmark polishing.

Related event: Anthropic Releases Claude Opus 5: SOTA Performance at Half the Price(128 posts)→

Original post →

More from Models

Models channel →