Claude Opus 5 Scores Three Times Higher Than Next Best Model on ARC-AGI-3
claudeai · x · 2026-07-25
Anthropic announced Claude Opus 5's performance on the ARC-AGI-3 evaluation, which tests AI models on solving novel problems. Opus 5 scored three times as high as the next best model, demonstrating exceptional reasoning and generalization capabilities.
Related event: Anthropic Unveils Claude Opus 5 with Top Benchmark Scores(28 posts)→
More from Models
- Vercel adds Claude Opus 5 to AI Gateway with fast mode for coding agents — EricBuess · 2026-07-25
- Early Claude Opus 5 feedback says it helps ship PRs faster in Claude Code — EricBuess · 2026-07-25
- Claude Opus 5 feels like Fable, but cheaper — cedric_chee · 2026-07-25
- Anthropic’s Opus 5 looks more like a major upgrade than a minor refresh — yi_ding · 2026-07-25
- Anthropic launches Claude Opus 5, with blind tests placing it above GPT-5.6 — lennysan · 2026-07-25
- Anthropic says Claude Opus 5 was intentionally left untrained on cyber tasks — rez0__ · 2026-07-25