New paper: error-ridden unsaturated benchmarks vastly underestimate AI capabilities

emollick · x · 2026-09-16

Ethan Mollick argues the state of public AI benchmarking is dire and undermining our ability to judge how good AI actually is. Citing a new paper, he notes most famous measures are maxed out (saturated), while the non-saturated benchmarks are riddled with so many errors that they vastly underestimate AI abilities.

Related event: Mollick: public AI benchmarks are broken and underestimate models(2 posts)→

Original post →

More from Research

Research channel →