Opus 5's High ARC-AGI-3 Score Sparks Cheating Allegations and Benchmark Debate

Claude Opus 5 achieved a score three times higher than its peers on the ARC-AGI-3 benchmark, but this result immediately sparked widespread skepticism and controversy within the AI community, with multiple users and industry insiders doubting its true capabilities.

Confirmed

It is currently confirmed that Opus 5's score has triggered a severe crisis of trust within the industry. Netizen @sdnr8 questioned the opacity of closed-source models, suggesting that the high score might not rely on "pure model" capabilities, but rather on external scaffolding like loops or harnesses. User @VraserX directly accused Anthropic of "cheating," alleging they specifically tailored their training to the puzzle patterns of this benchmark, converting visual reasoning into explicit algebraic problems for repetitive drilling. Furthermore, individuals claiming to be from frontier AI labs told @flowersslop that the score "looks fake," either due to excessive benchmark gaming or because competitors severely underestimated Anthropic's lead.

Unconfirmed

Whether Opus 5 actually utilized external agent frameworks, or if Anthropic indeed engaged in targeted specialized training, remains at the stage of community speculation and accusation. The relevant leaks have not provided substantive technical evidence.

Why It Matters

This controversy affects more than just Anthropic's reputation; it reflects deeper industry-wide concerns regarding the validity of AI evaluation systems. On one hand, as highlighted in a perspective shared by @morqon, evaluations that saturate too quickly lose their discriminative power, with some arguing that this benchmark was nearly "solved" on day one. On the other hand, discussions forwarded by @burny_tech indicate a growing consensus that ARC-AGI progress is primarily driven by targeted RL environments, prompting calls for ARC-AGI-4 to restrict public demos to prevent model overfitting. This demonstrates that existing public benchmarks are facing severe methodological challenges.

2026-07-25 ~ 2026-07-25 · 7 related posts

Full story(20 episodes)→

Primary sources