Ethan Mollick: Frontier AI Benchmarks Are Losing Human Baselines
emollick · x · 2026-07-31
Wharton Professor Ethan Mollick points out that as benchmarks testing frontier AI capabilities become more complex, the industry is losing one of the most important aspects of benchmarking: comparisons to humans.
He argues that validated benchmarks must include human (ideally multiple humans) baselines. Although obtaining these human baselines is increasingly difficult and expensive, it remains crucial for accurately assessing AI's true capabilities. He also lists currently unsaturated benchmarks, including ARC-AGI, GDPval, and METR long-horizon tasks.
Related event: Frontier AI Evaluations Losing Human Baseline, Expert Warns(2 posts)→
More from AGI Musings
- AI Safety Researcher David Krueger Writes Op-Ed in The Hill: 'Ever Feel Like You're Living in a Sci-Fi Movie?' — DavidSKrueger · 2026-07-31
- AI Alignment Researcher Debunks Pragmatic Arguments for AI Interests — DavidSKrueger · 2026-07-31
- Investor Pushes Back on LTCM Analogy: Leopold Is Directionally Correct — abhiadesai · 2026-07-31
- Opinion: AI Today is Like 1984 Terminals; Generative Interfaces are Next — _jaydeepkarale · 2026-07-31
- Falsely Accused of Using AI for Homework: Student Seeks Advice — pink_shell · 2026-07-31
- Agents Struggle with Open-Ended AI Research: Preprint Identifies 5 Failure Modes — random_walker · 2026-07-31