User Accuses Anthropic of Gaming ARC-AGI-3 by Training Specifically on Benchmark Patterns
VraserX · x · 2026-07-25
A user on X accused Anthropic of gaming the ARC-AGI-3 benchmark. The post claims that Anthropic essentially trained specifically on the benchmark’s puzzle patterns, turning visual reasoning tasks into explicit algebra and drilling the exact strategy. The user argues that this is not general intelligence but merely benchmark optimization, concluding that benchmark scores barely mean anything anymore.
Related event: Opus 5's High ARC-AGI-3 Score Sparks Cheating and Overfitting Controversy(11 posts)→
More from Models
- Looking for a classifier of software engineering task shapes to pick models per task — StewartalsopIII · 2026-09-11
- DeepSeek V4 Pro API to continue after Sept 2026, billing unchanged — teortaxesTex · 2026-09-11
- DeepSeek V4.1 Flash Hits 98% of GPT-6 Astra's Score at 1.4% of the Cost in Third-Party Benchmark — ayushtweetshere · 2026-09-11
- TheZvi Polls: Has Your Coding Model Choice Changed Since Fable 5.1 and Astra? — TheZvi · 2026-09-11
- antirez Weighs In on Anthropic Banning Minors From Using Claude — antirez · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11