Anthropic’s ARC-AGI-3 lead is being called meaningless as the benchmark saturates
morqon · x · 2026-07-25
- The thread pushes back on Anthropic’s claim that Claude Opus 5 is three times better than the next-best model on ARC-AGI-3, arguing that the benchmark was effectively solved on day one and may be of limited value.
- The cited post highlights a broader criticism of evals that are already saturated: once a benchmark is too easy, large score gaps stop saying much about real-world capability.
- The takeaway is less about one score and more about how quickly some public benchmarks become obsolete as models improve.
More from Models
- Moonshot's Kimi K3 Drops Monday; Baseten Offers Free API Credits — baseten · 2026-07-25
- Claude Opus 5 reportedly scores a perfect 42/42 on the 2026 IMO — exordin26 · 2026-07-25
- Claude Opus 5 launches with Box reporting big gains on enterprise agent tasks — inductionheads · 2026-07-25
- Critic says Gemini 3.5 Pro is already too late to compete — teortaxesTex · 2026-07-25
- Claude Opus 5 is now available in GitHub Copilot and Microsoft Foundry — DanWahlin · 2026-07-25
- Hyperagent says Opus 5 is stronger, but GPT-5.6 Sol is cheaper to deploy — TawohAwa · 2026-07-25