New paper: error-ridden unsaturated benchmarks vastly underestimate AI capabilities
emollick · x · 2026-09-16
Ethan Mollick argues the state of public AI benchmarking is dire and undermining our ability to judge how good AI actually is. Citing a new paper, he notes most famous measures are maxed out (saturated), while the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.
Related event: Mollick: public AI benchmarks are broken and underestimate models(2 posts)→
More from Research
- David Chalmers cites Jacques Thibodeau's work in his latest paper — JacquesThibs · 2026-09-17
- Insilico Medicine unveils 'Longevity Vaccines' combining AI and circular mRNA — rand_longevity · 2026-09-17
- NVIDIA's Axolotl3D does occlusion-aware 3D shape completion from multimodal inputs — NVIDIAAI · 2026-09-17
- alphaXiv maps every AI researcher's coauthorship graph to compute your LeCun number — burny_tech · 2026-09-17
- Gensyn releases open-1b, the first fully auditable and replayable foundation model — flngr · 2026-09-17
- Looped Transformer's Recurrent Depth: More Reasoning Without Extra Params or Longer CoT — gordic_aleksa · 2026-09-17