Anthropic’s Claude Opus 5 scores 30.2% on ARC-AGI-3, crushing the old 7.8% record
The Decoder · rss · 2026-07-26
According to The Decoder, Anthropic’s Claude Opus 5 scored 30.2% on ARC-AGI-3, a benchmark meant to measure general intelligence, nearly quadrupling the previous record of 7.8% set by GPT-5.6 Sol.
The benchmark’s developers said the model independently derived reflection equations, something they had not seen from other models before, and interpreted that as evidence of stronger logical reasoning.
More from Models
- Macaron-V1 lands on Hugging Face, built on Qwen3.6-35B-A3B — MundanePercentage674 · 2026-07-26
- POCKET-35B claims 59 tok/s on CPU and runs locally on phones without a GPU — Powerful_Evening5495 · 2026-07-26
- Runway Agent still tops the author’s test of major video agents — taherdhanera · 2026-07-26
- SemiAnalysis: Kimi K3 Beats NVIDIA's Nemotron, Calls Committee Model a Failure — FinanceYF5 · 2026-07-26
- Users report Claude outages again as status page still shows mostly healthy service — Ubunta · 2026-07-26
- Google AI Overview turns a joke search into a mythic warning about Polyphemus — ctjlewis · 2026-07-26