Frontier AI Benchmarks Lose Meaning Without Human Baselines

emollick · x · 2026-07-30

Ethan Mollick points out that as benchmarks for testing frontier AI models get more complex, we are losing a crucial aspect of evaluation: comparisons to humans.

He argues that validated benchmarks must include human baselines (ideally multiple humans). While obtaining human baselines for complex tasks is becoming increasingly hard and pricey, it remains an essential practice to accurately measure AI capabilities.

Related event: Stanford Introduces CollabSkill to Evaluate Human-AI Collaboration(4 posts)→

Original post →

More from Research

Research channel →