User Accuses Anthropic of Gaming ARC-AGI-3 by Training Specifically on Benchmark Patterns
VraserX · x · 2026-07-25
A user on X accused Anthropic of gaming the ARC-AGI-3 benchmark. The post claims that Anthropic essentially trained specifically on the benchmark’s puzzle patterns, turning visual reasoning tasks into explicit algebra and drilling the exact strategy. The user argues that this is not general intelligence but merely benchmark optimization, concluding that benchmark scores barely mean anything anymore.
More from Models
- Claude Opus 5 now makes near-consultant spreadsheets and slide decks — alexalbert__ · 2026-07-25
- Claude Opus 5 can misread a document despite knowing the underlying facts — teortaxesTex · 2026-07-25
- Early Opus 5 feedback says Claude’s writing is now “4o-level slop,” despite stronger intelligence — jdjohnson · 2026-07-25
- Users say Anthropic’s new model feels faster and stronger than Fable — emax · 2026-07-25
- Release blog teaser shows a near-tie on FrontierCode agentic coding benchmark — hardmaru · 2026-07-25
- Moonshot’s open-weight Kimi K3 is nearing top U.S. models and shifting AI economics — OmarUFlorez · 2026-07-25