Claude Opus 5 Sets New ARC-AGI-3 Record

Claude Opus 5 achieved a score of 30.2% in the highly challenging public ARC-AGI-3 demo, becoming the first model to demonstrate clearly usable performance. It massively broke the previous record of 7.8% held by GPT-5.6 Sol (Max) and showcased a novel ability to translate visual puzzles into algebraic notation for reasoning, drawing wide attention from the AI community.

Confirmed

Based on information from ARC Prize and analysis of Claude Opus 5's system card, the following facts are confirmed:

Unconfirmed

Why it matters

The ARC-AGI benchmark is notoriously brutal, designed to test generalization on entirely novel tasks. Claude Opus 5's leap in absolute score and its emergent algebraic reasoning strategy indicate a potential new mechanism for abstract logic and complex problem-solving, serving as a crucial metric for evaluating the intelligence of future frontier models.

2026-07-25 ~ 2026-07-27 · 15 related posts

Full story(14 episodes)→

Primary sources

3 near-duplicate retellings: GregKamradt · EricBuess · rbhar90