ARC-AGI-3 Score Jumps to 96%: Giving Opus 5 a Computer to Build Its Own Tools

新智元 · wechat · 2026-08-13

A test by developer Jeremy Berman has sparked debate: by changing the runtime environment, Claude Opus 5's score on the ARC-AGI-3 benchmark surged from the official 30.2% to 96.2%.

The breakthrough lies in 'giving the model a computer': The test used no elaborate prompts or specific code, only providing a Claude Code environment, an action command, and a file system log. Faced with unfamiliar game levels, Opus 5 autonomously figured out the rules and wrote parsers, search functions, and even game simulators on the fly (totaling 12,700 lines of code), discarding them after clearing the level.

Results & Reflections:

Related event: Opus 5 with Claude Code Scores 96% on ARC-AGI-3 Benchmark(3 posts)→

Original post →

More from coding & agent

coding & agent channel →