Claude Opus 5 hits 30.2% on ARC-AGI-3, far ahead of GPT-5.6 Sol
rbhar90 · x · 2026-07-25
Claude Opus 5 is described as the first model to do meaningfully well on ARC-AGI-3, reaching 30.2% on the benchmark.
- The previous best score was 7.8%, set by GPT-5.6 Sol (Max).
- In a replay review, the poster notes that on the semi-private set Opus 5 fully solves about 20% of games, but scores 0 on another 20%, showing a stark easy-vs-hard split.
- The observed successes are attributed to effective goal recognition, hypothesis exploration, and context carry-forward, which the poster argues may be a step toward agentic intelligence in novel environments.
- A linked note from ARC says Opus 5 shows novel behavior and outperforms Fable on previously unbeaten environments.
Related event: Claude Opus 5 Sets New SOTA on ARC-AGI-3 with Algebraic Reasoning(7 posts)→
More from Models
- PaddlePaddle’s HPD-Parsing trends on Hugging Face as a document parsing model — PaddlePaddle · 2026-07-25
- Claude Opus 5 peaks at high reasoning level on Vibe Code Bench, then gets pricier — thesaraharminta · 2026-07-25
- Gemini Pro tells a user to search the web, then says it can’t access the content — gaganghotra_ · 2026-07-25
- Matt Shumer teases an Opus 5 demo that could surprise AI Twitter tomorrow — mattshumer_ · 2026-07-25
- Meme mocks the exploding AI model release maze — kevinnbass · 2026-07-25
- Qwen3.5-122B runs on OrangePi AI Studio Pro after a device-capability shim fixes vLLM — StillVeterinarian578 · 2026-07-25