Mollick: public AI benchmarks are broken and underestimate AI abilities

emollick · x · 2026-09-16

Ethan Mollick argues the state of public AI benchmarking is "dire" and undermining our ability to gauge current AI. Per the cited paper, most famous benchmarks are saturated, while the non-saturated ones are riddled with errors that vastly underestimate AI abilities.

Related event: Mollick: public AI benchmarks are broken and underestimate models(2 posts)→

Original post →

More from Models

Models channel →