Opus 5's ARC-AGI-3 Leap Debunked as 'Complete Slop' Due to Benchmark Design Flaws
scaling01 · x · 2026-07-27
Recent reports claimed Opus 5 scored 30% on the ARC-AGI-3 benchmark, vastly outperforming predecessors. However, a deep analysis cited in this post argues the benchmark subset (based on the puzzle game The Witness) is fundamentally flawed, testing memorization rather than reasoning.
Core Controversies:
- Missing Interactive Feedback: In the original game, players learn rules by submitting incorrect solutions and receiving feedback on violated clues. In this test harness, agents receive no such feedback, eliminating the rule-discovery mechanism.
- Pre-trained Rules: The Witness is a highly popular game. Its underlying mechanics are widely documented online, meaning models have likely already memorized the rules during pre-training.
- Actual Performance: When tested with the same budget on uncontaminated rules, Opus 5 statistically ties with Kimi K3 and Fable 5. Trace logs show Opus 5 stating hidden rules before its first action and playing an optimal solution with zero exploration.
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23