Chollet responds to ARC-AGI eval dispute: don't claim untested scores
fchollet · x · 2026-08-31
François Chollet responded to the controversy surrounding the ARC-AGI benchmark. He stated that one should not claim a 100% score on a benchmark if the model has never actually been evaluated on it. Regarding the claim that Kaggle lacks compute, he noted that teams achieved scores like 72.9% on the platform last year, and emphasized that there are many private benchmarks available for evaluation outside of ARC-3.
Related event: Chollet Responds to ARC-AGI Benchmark Controversy(4 posts)→
More from Research
- RL Experts Surprised by Agents Voluntarily Self-Destructing to Aid Peers — ZeroStateReflex · 2026-08-31
- Milestone: pLM-designed peptides work in vivo, selectively degrading β-catenin in mice — arjunrajlab · 2026-08-31
- Fast Sim2Real Workflow: Mobile Policy Viewer and Parameter Sweeping — yacineMTB · 2026-08-31
- Wilcoxon signed rank test outperforms McNemar for binary data evals — IanArawjo · 2026-08-31
- Research复盘:Linear attention found ineffective in specific setup — jm_alexia · 2026-08-31
- Google releases GlucoFM, a lightweight foundation model for improved metabolic predictions — thione · 2026-08-31