Claude Opus 5 scores three times higher than the runner-up on ARC-AGI-3
claudeai · x · 2026-07-25
Anthropic says Claude Opus 5 reaches three times the score of the next best model on ARC-AGI-3, an evaluation built around novel problem solving.
The accompanying chart also frames the result by cost, showing Opus 5 as a strong performer at roughly comparable or lower per-task cost than rivals.
Related event: Anthropic Releases Claude Opus 5 with New SOTA Performance at Half Price(25 posts)→
More from Models
- Claude Opus 5 feels like Fable, but cheaper — cedric_chee · 2026-07-25
- Anthropic’s Opus 5 looks more like a major upgrade than a minor refresh — yi_ding · 2026-07-25
- Anthropic launches Claude Opus 5, with blind tests placing it above GPT-5.6 — lennysan · 2026-07-25
- Anthropic’s Claude Opus 5 scores 43.3% on Frontier-Bench and stays at $5/$25 — mark_k · 2026-07-25
- Claude Opus 5’s ARC-AGI-3 chart drew a one-word reaction — var_epsilon · 2026-07-25
- Notion acquires ZeroEntropy and its reranker ships open source under Apache 2.0 — CShorten30 · 2026-07-25