Mollick: public AI benchmarks are broken and underestimate models

Ethan Mollick shared a new paper arguing that public AI benchmarks are in bad shape: most famous ones are saturated and the unsaturated ones are riddled with errors, causing systematic underestimation of AI capabilities.

2026-09-16 ~ 2026-09-16 · 2 related posts