Opus 5 scores 96.2% on ARC-AGI-3 with generic agent, near-human performance

inductionheads · x · 2026-08-14

Developer Jeremy Berman announced on X that using Claude Code with Opus 5 (high) achieved 96.2% accuracy on the ARC-AGI-3 benchmark, with pass@2 at 99.3%. The program is almost entirely generic, relying on agent capabilities rather than ARC-specific tweaks. Plinz commented that François Chollet's approach of challenging researchers to prove him wrong is productive and courageous.

Related event: Opus 5 with Claude Code Scores 96% on ARC-AGI-3 Benchmark(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →