Stanford studies find AI benchmarks may not measure what they claim, with billions riding on scores

StanfordHAI · x · 2026-10-06

Two new Stanford-supported studies, to be presented at the third COLM conference in San Francisco this October, turn the spotlight on AI benchmarks themselves: the standardized tests that score models on reasoning, safety and bias drive billions in funding and increasingly shape regulation and government procurement — yet they may not measure what they claim to.

Assistant professor Sanmi Koyejo and graduate student Sang Truong argue the field needs to approach benchmark construction with far more rigor. Key points:

Original post →

More from Safety

Safety channel →