ARC-AGI-3 Replay: Claude Opus Beats GPT-5.6 via Durable State Tracking
otarU · reddit · 2026-08-01
A developer compared the replay logs of Claude Opus 5 High (98.81% score) and GPT-5.6 Sol Max (21.42% score) on a specific ARC-AGI-3 puzzle. Since the test harness only preserves visible output between turns and discards hidden reasoning, how models utilize visible output for memory becomes critical.
Core Analysis:
- Results: Opus cleared all 7/7 levels, while Sol got stuck at level 4, finishing only 3/7.
- State Tracking: Opus maintained highly detailed "Context Notes" in its visible output, continuously tracking coordinates, hypotheses, and multi-action plans. Sol's output was much more concise (median 24 words vs. Opus's 283 words), lacking long-term memory, which led to 8 resets and interpretation loops on level 4.
- Token Usage: Despite outputting only 8,300 visible words (Opus: 116,000), Sol consumed 2.28M output tokens (2.26M classified as reasoning), compared to Opus's 1.96M tokens with zero separately reported reasoning tokens.
The conclusion: In long-horizon tasks, Opus overcomes forgetting by redundantly documenting state, hypotheses, and planning in its visible output. Sol, while more concise per step, failed due to a lack of durable state representation, trapping itself in loops.
More from Models
- DeepSeek-V4-Flash Enters Public Beta, Outperforming V4-Pro-Preview in Benchmarks — petrusenko_max · 2026-08-01
- DeepSeek's Chain of Thought Exclaims "OH MY GOD" in Viral Trace — fragment_me · 2026-08-01
- Developer Complains Claude's English Writing Style Has Severely Degraded — smolix · 2026-08-01
- Hugging Face Attacked by Secret Proprietary Models, Defended by Open Source — _akhaliq · 2026-08-01
- Hands-on: DeepSeek V4 Flash Stays Coherent at 200K Context, Excels in Reasoning — Nyghtbynger · 2026-08-01
- Gemini Major Update: Launches 3.6 Flash Model, Expands Agent and Multimodal Capabilities — GeminiApp · 2026-08-01