FULL STORY

Claude Opus Faces ARC-AGI Evaluation Controversy

Developers questioned Claude Opus's high ARC-AGI score due to API implementation issues, prompting the official team to clarify the testing standards and open-source the ARC-AGI-3 benchmark.

2026-07-30 ~ 2026-07-30 · 2 episodes · 6 posts

Episode 1 · Claude Opus ARC-AGI Score Questioned Over API Flaw (2026-07-30, 2 posts)

A developer argued that Claude Opus's high ARC-AGI score was inflated by an API flaw, but community members maintain that its generalization capabilities still outperform GPT-5 under the same testing framework.

Episode 2 · ARC-AGI-3 Officially Open-Sources Benchmarking Codebase (2026-07-30, 4 posts)

Greg Kamradt clarified that identical sliding windows were used for OpenAI and Anthropic in the ARC-AGI evaluation, addressing recent cheating concerns. Additionally, ARC Prize officially open-sourced the ARC-AGI-3 benchmarking codebase to help evaluate complex reasoning in frontier models.