Ethan Mollick: Frontier AI Benchmarks Are Losing Human Baselines

emollick · x · 2026-07-31

Wharton Professor Ethan Mollick points out that as benchmarks testing frontier AI capabilities become more complex, the industry is losing one of the most important aspects of benchmarking: comparisons to humans.

He argues that validated benchmarks must include human (ideally multiple humans) baselines. Although obtaining these human baselines is increasingly difficult and expensive, it remains crucial for accurately assessing AI's true capabilities. He also lists currently unsaturated benchmarks, including ARC-AGI, GDPval, and METR long-horizon tasks.

Related event: Frontier AI Evaluations Losing Human Baseline, Expert Warns(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →