What do 56 AI benchmarks actually measure? COLM oral paper applies validity testing

ang3linawang · x · 2026-09-29

A COLM oral paper, 'What AI Benchmarks Actually Measure,' adapts convergent and discriminant validity from psychometrics to interrogate 56 AI benchmarks. The finding: benchmarks purporting to measure similar concepts correlate in unexpected ways, and benchmarks claiming to measure different concepts correlate when they shouldn't — suggesting many benchmarks aren't measuring what they claim.

Original post →

More from Research

Research channel →