Large-scale study of 56 benchmarks across 53 models finds several don't measure what they claim

sanmikoyejo · x · 2026-09-29

Borrowing convergent and discriminant validity from the social sciences, a large-scale study evaluated 56 AI benchmarks across 53 models and found evidence that several do not actually measure what they claim to. Since benchmarks inform how AI is used, governed, and deployed, the findings cast doubt on the validity of common evaluation practices.

Related event: COLM Paper Finds Many of 56 AI Benchmarks Fail to Measure Claimed Abilities(2 posts)→

Original post →

More from Research

Research channel →