Opus 5's High ARC-AGI-3 Score Sparks Cheating and Overfitting Controversy

Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, outperforming peers by 3 times. However, this high score immediately triggered widespread allegations of cheating and overfitting within the AI community, with multiple users and industry insiders questioning its true capabilities and sparking reflections on the validity of evaluation systems and their ties to commercial interests.

Confirmed

It is confirmed that Opus 5's specific score is 30.2%, and this result has caused a severe crisis of confidence in the industry. Reddit user @sdnr8 questioned the opacity of closed-source models, suggesting the high score might not come from "pure model" capabilities but from external proxy frameworks like loop or harness. User @VraserX directly accused Anthropic of "cheating," claiming they specifically trained for the benchmark's puzzle patterns, converting visual reasoning into explicit algebraic problems. Furthermore, a person claiming to be from a frontier AI lab told @flowersslop that the score "looks fake," implying either excessive benchmark gaming or a severe underestimation of Anthropic's lead by competitors. Chris Szegedy (@ChrSzegedy) and @DrSingularity explicitly stated that ARC-AGI has little to do with AGI, though the latter acknowledged that AI is indeed getting stronger. Gary Marcus also emphasized that scoring high on benchmarks does not equate to approaching AGI, arguing that the score improvement is more likely due to targeted optimization rather than the true generalization of abstract reasoning capabilities.

Unconfirmed

Whether Opus 5 actually used external proxy frameworks, or whether Anthropic indeed conducted targeted specialized training, currently remains at the stage of community speculation and accusations, with no substantial technical evidence provided for the relevant claims.

Why it matters

This controversy affects more than just Anthropic's reputation; it reflects the industry's deep-seated concerns about the validity of AI evaluation systems. On one hand, as relayed by @morqon, rapidly saturating evaluations lose their discriminative power, with some believing the benchmark was nearly "solved" on day one. On the other hand, discussions forwarded by @burnytech indicate a growing default assumption that ARC-AGI progress mainly comes from targeted RL environments, prompting calls for ARC-AGI-4 to reduce public demos to prevent model overfitting. @JasonBotteril also pointed out that naming a benchmark directly after AGI is almost destined to induce labs to optimize for it specifically. In addition, @rbhar90 emphasized that the "strong reasoning ability" narrative of frontier labs is deeply bound to huge commercial interests, advocating for more open, verifiable benchmarks (such as the self-built Witness test set), and noting that Opus 5's improvement on ARC-AGI-3 did not transfer to other test sets. @inductionheads also noted that if a model is indeed trained in an RL environment similar to ARC-AGI, its performance cannot prove "generalization." This indicates that existing public benchmarks are facing severe methodological challenges.

2026-07-25 ~ 2026-07-27 · 11 related posts

Full story(9 episodes)→

Primary sources