Opus 5's ARC-AGI-3 Leap Debunked as 'Complete Slop' Due to Benchmark Design Flaws
scaling01 · x · 2026-07-27
Recent reports claimed Opus 5 scored 30% on the ARC-AGI-3 benchmark, vastly outperforming predecessors. However, a deep analysis cited in this post argues the benchmark subset (based on the puzzle game The Witness) is fundamentally flawed, testing memorization rather than reasoning.
Core Controversies:
- Missing Interactive Feedback: In the original game, players learn rules by submitting incorrect solutions and receiving feedback on violated clues. In this test harness, agents receive no such feedback, eliminating the rule-discovery mechanism.
- Pre-trained Rules: The Witness is a highly popular game. Its underlying mechanics are widely documented online, meaning models have likely already memorized the rules during pre-training.
- Actual Performance: When tested with the same budget on uncontaminated rules, Opus 5 statistically ties with Kimi K3 and Fable 5. Trace logs show Opus 5 stating hidden rules before its first action and playing an optimal solution with zero exploration.
More from Models
- Google’s Gemini 3.6 Flash could win by matching Sonnet quality at a lower cost — haider1 · 2026-07-27
- Users Report Severe Downgrade in Opus 5: Hallucinations and Math Errors — whatsallthiss · 2026-07-27
- A Reddit explainer breaks down MoE, KV cache, MLA and KDA behind Kimi K3 — MohamedKadri_ · 2026-07-27
- Reddit asks which local model works best for coding, planning and VS Code workflows — naunen · 2026-07-27
- GPT-5.6 Sol fixes 31 of 105 hidden bugs in a two-repo benchmark — PawelHuryn · 2026-07-27
- MineBench says Opus 5.0 beats Fable 5 on build quality but costs 64% more — ENT_Alam · 2026-07-27