Chollet debunks "100% on ARC-AGI-3" claim: only on easy public set

fchollet · x · 2026-09-01

François Chollet responded to a claim about a coding agent scoring 100% on ARC-AGI-3. He pointed out that the ARC-AGI-3 paper stated in March that the public evaluation set is "intentionally easier for both humans and AI" for demonstration purposes. He rebutted the defense that the public set is a valid leaderboard, arguing that scoring high on these simple games doesn't equate to mastering the task, and criticized the dismissal of the lack of semi-private data validation.

Related event: Chollet Pushes Back on Agent Claiming 100% on ARC-AGI-3(6 posts)→

Original post →

More from Models

Models channel →