Opus 5 with Claude Code Scores 96% on ARC-AGI-3 Benchmark
By changing the runtime environment to give the model computer access, developer Jeremy Berman boosted Claude Opus 5's score on the ARC-AGI-3 benchmark from 30.2% to 96.2%. Using Claude Code with minimal specialized code, the setup nearly matched human-level performance.
2026-08-13 ~ 2026-08-14 · 3 related posts
- Claude Code Hits 96.2% on ARC-AGI-3 with Almost Zero ARC-Specific Code — khademinori · 2026-08-13
- ARC-AGI-3 Score Jumps to 96%: Giving Opus 5 a Computer to Build Its Own Tools — 新智元 · 2026-08-13
- Opus 5 scores 96.2% on ARC-AGI-3 with generic agent, near-human performance — inductionheads · 2026-08-14