Anthropic is using ARC-AGI-3 to gauge Opus 5’s novel problem-solving ability

typewriters · x · 2026-07-25

Anthropic is reportedly using ARC-AGI-3 to measure the novel problem-solving ability of Opus 5.

The attached chart compares models by total evaluation cost on the x-axis and score on the y-axis. In the visual, Opus 5 (high) is shown around 30% on ARC-AGI-3, while Opus 4.8 (high) sits near the low single digits and GPT-5.6 Sol is plotted below 10% at higher cost.

The post frames ARC-AGI-3 as a benchmark for frontier-style novel problem solving rather than routine benchmark polishing.

Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(32 posts)→

Original post →

More from Models

Models channel →