Chollet defends ARC-AGI-3: benchmark well calibrated, smart humans can score over 90%

fchollet · x · 2026-09-04

François Chollet pushes back on critics of ARC-AGI-3. When it launched in March, frontier models scored under 1%, and some Singularitarian posters attacked the benchmark as broken — claiming even the smartest humans couldn't solve it or that the max score was 40%. Chollet says the benchmark is perfectly calibrated: a smart human putting in real effort can score above 90%, and 100% is achievable by simply beating the (weak, unfiltered) human baseline on action count. By symmetry, AI can also reach 100% once real progress toward agentic general intelligence is made.

Original post →

More from Models

Models channel →