ARC-AGI-3 Officially Open-Sources Benchmarking Codebase
Greg Kamradt clarified that identical sliding windows were used for OpenAI and Anthropic in the ARC-AGI evaluation, addressing recent cheating concerns. Additionally, ARC Prize officially open-sourced the ARC-AGI-3 benchmarking codebase to help evaluate complex reasoning in frontier models.
2026-07-30 ~ 2026-07-30 · 4 related posts
- Episode 1: Claude Opus ARC-AGI Score Questioned Over API Flaw(2026-07-30, 2 posts)
- Episode 2: ARC-AGI-3 Officially Open-Sources Benchmarking Codebase(2026-07-30, 4 posts)
- Greg Kamradt Responds to Evaluation Dispute: Same Rolling Window Used for Opus and OpenAI — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repo Goes Open Source — GregKamradt · 2026-07-30
- ARC-AGI-3 Official Benchmarking Repository Open-Sourced — GregKamradt · 2026-07-30
- ARC-AGI-3 API clarification: 'reasoning' field logs model output, not private CoTs — GregKamradt · 2026-07-30