FULL STORY
Claude Opus Faces ARC-AGI Evaluation Controversy
Developers questioned Claude Opus's high ARC-AGI score due to API implementation issues, prompting the official team to clarify the testing standards and open-source the ARC-AGI-3 benchmark.
2026-07-30 ~ 2026-07-30 · 2 episodes · 6 posts
Episode 1 · Claude Opus ARC-AGI Score Questioned Over API Flaw (2026-07-30, 2 posts)
A developer argued that Claude Opus's high ARC-AGI score was inflated by an API flaw, but community members maintain that its generalization capabilities still outperform GPT-5 under the same testing framework.
- Netizen Questions: If Harness is Identical, Opus Generalizes Better Than GPT-5 — umike_njsf · 2026-07-30
- ARC-AGI Evaluation Dispute: Claude Opus Score Questioned Over API Implementation Flaw — steipete · 2026-07-30
Episode 2 · ARC-AGI-3 Officially Open-Sources Benchmarking Codebase (2026-07-30, 4 posts)
Greg Kamradt clarified that identical sliding windows were used for OpenAI and Anthropic in the ARC-AGI evaluation, addressing recent cheating concerns. Additionally, ARC Prize officially open-sourced the ARC-AGI-3 benchmarking codebase to help evaluate complex reasoning in frontier models.
- Greg Kamradt Responds to Evaluation Dispute: Same Rolling Window Used for Opus and OpenAI — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repo Goes Open Source — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repository Open-Sourced — GregKamradt · 2026-07-30
- ARC-AGI-3 API clarification: 'reasoning' field logs model output, not private CoTs — GregKamradt · 2026-07-30