What do 56 AI benchmarks actually measure? COLM oral paper applies validity testing
ang3linawang · x · 2026-09-29
A COLM oral paper, 'What AI Benchmarks Actually Measure,' adapts convergent and discriminant validity from psychometrics to interrogate 56 AI benchmarks. The finding: benchmarks purporting to measure similar concepts correlate in unexpected ways, and benchmarks claiming to measure different concepts correlate when they shouldn't — suggesting many benchmarks aren't measuring what they claim.
More from Research
- 30+ labs fail to replicate Marcus et al's 1999 Science paper on infants vs RNNs, built on just 16 babies — tallinzen · 2026-09-29
- Chris Manning on why LMs learn verb categories first, and why he thinks LeCun is wrong about language — ziv_ravid · 2026-09-29
- Researchers clash over what a valid Bayesian updating rule in research synthesis should look like — RexDouglass · 2026-09-29
- Google Research unveils multi-agent 'AI video co-director' for consistent long-form video generation — rseroter · 2026-09-29
- BLUE from Google internship lands NeurIPS: LLM-written user profiles boost recommendations — shangbinfeng · 2026-09-29
- Collatz twist: Krasikov-Lagarias-style X^0.84 bounds apply to any root, and equally to 3x-1 — AlexKontorovich · 2026-09-29