Opus 5 appears to improve on ARC-AGI 1 and 2, and may rely on algebraic puzzle solving
herbiebradley · x · 2026-07-25
Herbie Bradley says the Opus 5 system card shows clear gains on ARC-AGI 1 and 2 compared with 4.7/4.8.
He also notes that the system-card description of an ARC-AGI-3-like task sounds like the model learns to turn a visual puzzle into explicit algebra, which he sees as a very trainable strategy. Based on that, he predicts the improvement likely came from a focused research project on these puzzle types, possibly using a data vendor under an exclusivity agreement.
More from Models
- AutomationBench chart puts Opus 5 ahead on long-horizon agent tasks — daniel_mac8 · 2026-07-25
- Claude Opus 5 is now available in GitHub Copilot and Microsoft Foundry — DanWahlin · 2026-07-25
- CursorBench 3.2 puts Claude Opus 5 within 0.5 points of Fable 5 at half the task cost — EricBuess · 2026-07-25
- Hyperagent says Opus 5 is stronger, but GPT-5.6 Sol is cheaper to deploy — TawohAwa · 2026-07-25
- Opus 5 Reportedly Crushes Fable 5 in Benchmarks as Model Wars Heat Up — haider1 · 2026-07-25
- Qwen3.5-9B uncensored GGUF variant starts trending on Hugging Face — DavidAU · 2026-07-25