User Accuses Anthropic of Gaming ARC-AGI-3 by Training Specifically on Benchmark Patterns

VraserX · x · 2026-07-25

A user on X accused Anthropic of gaming the ARC-AGI-3 benchmark. The post claims that Anthropic essentially trained specifically on the benchmark’s puzzle patterns, turning visual reasoning tasks into explicit algebra and drilling the exact strategy. The user argues that this is not general intelligence but merely benchmark optimization, concluding that benchmark scores barely mean anything anymore.

Original post →

More from Models

Models channel →