Claude Opus 5 hits 30.2% on ARC-AGI-3, far ahead of GPT-5.6 Sol
rbhar90 · x · 2026-07-25
Claude Opus 5 is described as the first model to do meaningfully well on ARC-AGI-3, reaching 30.2% on the benchmark.
- The previous best score was 7.8%, set by GPT-5.6 Sol (Max).
- In a replay review, the poster notes that on the semi-private set Opus 5 fully solves about 20% of games, but scores 0 on another 20%, showing a stark easy-vs-hard split.
- The observed successes are attributed to effective goal recognition, hypothesis exploration, and context carry-forward, which the poster argues may be a step toward agentic intelligence in novel environments.
- A linked note from ARC says Opus 5 shows novel behavior and outperforms Fable on previously unbeaten environments.
Related event: Claude Opus 5 Sets New ARC-AGI-3 Record(15 posts)→
More from Models
- BullshitBench update: GPT-6-Astra beats all prior OpenAI models but still trails Anthropic — scaling01 · 2026-09-11
- Astra Scores 83% on GauntletBench, First Computer-Use Agent to Beat Human Baseline — ducha_aiki · 2026-09-11
- Kimi K2.8 Preview rolls out: near-K3 coding performance, 1M context for all tiers — teortaxesTex · 2026-09-11
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11