Mollick: public AI benchmarks are broken and underestimate models
Ethan Mollick shared a new paper arguing that public AI benchmarks are in bad shape: most famous ones are saturated and the unsaturated ones are riddled with errors, causing systematic underestimation of AI capabilities.
2026-09-16 ~ 2026-09-16 · 2 related posts
- Mollick: public AI benchmarks are broken and underestimate AI abilities — emollick · 2026-09-16
- New paper: error-ridden unsaturated benchmarks vastly underestimate AI capabilities — emollick · 2026-09-16