Opus 5's ARC-AGI-3 Leap Debunked as 'Complete Slop' Due to Benchmark Design Flaws

scaling01 · x · 2026-07-27

Recent reports claimed Opus 5 scored 30% on the ARC-AGI-3 benchmark, vastly outperforming predecessors. However, a deep analysis cited in this post argues the benchmark subset (based on the puzzle game The Witness) is fundamentally flawed, testing memorization rather than reasoning.

Core Controversies:

Original post →

More from Models

Models channel →