ARC-AGI-3 Replay: Claude Opus Beats GPT-5.6 via Durable State Tracking

otarU · reddit · 2026-08-01

A developer compared the replay logs of Claude Opus 5 High (98.81% score) and GPT-5.6 Sol Max (21.42% score) on a specific ARC-AGI-3 puzzle. Since the test harness only preserves visible output between turns and discards hidden reasoning, how models utilize visible output for memory becomes critical.

Core Analysis:

The conclusion: In long-horizon tasks, Opus overcomes forgetting by redundantly documenting state, hypotheses, and planning in its visible output. Sol, while more concise per step, failed due to a lack of durable state representation, trapping itself in loops.

Original post →

More from Models

Models channel →