Opus 5 scores 96.2% on ARC-AGI-3 with generic agent, near-human performance
inductionheads · x · 2026-08-14
Developer Jeremy Berman announced on X that using Claude Code with Opus 5 (high) achieved 96.2% accuracy on the ARC-AGI-3 benchmark, with pass@2 at 99.3%. The program is almost entirely generic, relying on agent capabilities rather than ARC-specific tweaks. Plinz commented that François Chollet's approach of challenging researchers to prove him wrong is productive and courageous.
Related event: Opus 5 with Claude Code Scores 96% on ARC-AGI-3 Benchmark(3 posts)→
More from AGI Musings
- Debate Sparks: Yudkowsky Never Posed a Real ASI Danger, Researcher Argues — jd_pressman · 2026-08-14
- The Real Divide in AI Coding: Unsupervised Generation vs Verified Engineering — OGMYT · 2026-08-14
- xAI Exec Reiterates Open Source Value: Preventing Frontier Model Power Overconcentration — TinfoilTricorn · 2026-08-14
- AI to Reshape Construction: Today's Data Center Designers Will Transform Housing by 2032 — deanwball · 2026-08-14
- New Paper Explains How LLMs 'Hijack' Pleistocene Brains into Perceiving False Agency — MacrinePhD · 2026-08-14
- The Enterprise AI Trap: Why Most Companies Shouldn't Post-Train Models — xennygrimmato_ · 2026-08-14