Chollet defends ARC-AGI-3: benchmark well calibrated, smart humans can score over 90%
fchollet · x · 2026-09-04
François Chollet pushes back on critics of ARC-AGI-3. When it launched in March, frontier models scored under 1%, and some Singularitarian posters attacked the benchmark as broken — claiming even the smartest humans couldn't solve it or that the max score was 40%. Chollet says the benchmark is perfectly calibrated: a smart human putting in real effort can score above 90%, and 100% is achievable by simply beating the (weak, unfiltered) human baseline on action count. By symmetry, AI can also reach 100% once real progress toward agentic general intelligence is made.
More from Models
- Prompting Fable 5.1 with embodied mannerisms like *shrugs* makes roleplay flow better — repligate · 2026-09-05
- Astra appears in ChatGPT/Codex on Windows but not Mac, same account — cheesecakegood · 2026-09-05
- No usage reset on GPT-6 Astra launch day, developer calls out OpenAI — tobowers · 2026-09-05
- Debate: Models Fuzzily Recall Concepts, Not Text — SAE Features vs Edit-Distance Memorization — voooooogel · 2026-09-05
- Blogger feeds GPT6 Astra a PPT template, gets 30 conference slides with the right avatar — vista8 · 2026-09-05
- Dozens of GPT-6 Astra prompts collected via GPT 6 pro web search — vista8 · 2026-09-05