Anthropic’s Claude Opus 5 scores 30.2% on ARC-AGI-3, crushing the old 7.8% record

The Decoder · rss · 2026-07-26

According to The Decoder, Anthropic’s Claude Opus 5 scored 30.2% on ARC-AGI-3, a benchmark meant to measure general intelligence, nearly quadrupling the previous record of 7.8% set by GPT-5.6 Sol.

The benchmark’s developers said the model independently derived reflection equations, something they had not seen from other models before, and interpreted that as evidence of stronger logical reasoning.

Original post →

More from Models

Models channel →