Chollet debunks "100% on ARC-AGI-3" claim: only on easy public set
fchollet · x · 2026-09-01
François Chollet responded to a claim about a coding agent scoring 100% on ARC-AGI-3. He pointed out that the ARC-AGI-3 paper stated in March that the public evaluation set is "intentionally easier for both humans and AI" for demonstration purposes. He rebutted the defense that the public set is a valid leaderboard, arguing that scoring high on these simple games doesn't equate to mastering the task, and criticized the dismissal of the lack of semi-private data validation.
Related event: Chollet Pushes Back on Agent Claiming 100% on ARC-AGI-3(6 posts)→
More from Models
- NVIDIA reportedly buys Hugging Face for $12.9B; GLM-5.3-Flash and Qwen4 preview land — altryne · 2026-09-01
- METR + Redwood GPT agent swarm probe could only analyze GPT variants, not Claude — geoffreyirving · 2026-09-01
- A nonprofit wants Qwen for everyday paperwork: how to pick local hardware — DerAndi_DE · 2026-09-01
- Thomson Reuters builds $40M legal LLM rivaling Claude Opus on Qwen — josh_wills · 2026-09-01
- Are LLMs just statistical search engines? A rebuttal on why search calls aren't fundamental — gandamu_ml · 2026-09-01
- Teknium: Hermes has long supported NVIDIA's openshell — corporate IT teams that missed it are uninformed — Teknium · 2026-09-01