Claude Code + Opus 5 saturates ARC-AGI-3 via falsifiable predictions

scaling01 · x · 2026-08-20

A developer achieved 100% RHAE on the ARC-AGI-3 benchmark using a general-purpose coding agent (Claude Code + Opus 5) without additional harness.

The key idea was simple: force a falsifiable prediction before every action.

This turns every move into an experiment and every miss into a precise correction, allowing the agent to learn the game mechanics by being wrong on the record, effectively saturating the benchmark.

Related event: Claude Opus 5 with Claude Code Achieves Perfect Score on ARC-AGI 3(2 posts)→

Original post →

More from coding & agent

coding & agent channel →