Large-scale study of 56 benchmarks across 53 models finds several don't measure what they claim
sanmikoyejo · x · 2026-09-29
Borrowing convergent and discriminant validity from the social sciences, a large-scale study evaluated 56 AI benchmarks across 53 models and found evidence that several do not actually measure what they claim to. Since benchmarks inform how AI is used, governed, and deployed, the findings cast doubt on the validity of common evaluation practices.
Related event: COLM Paper Finds Many of 56 AI Benchmarks Fail to Measure Claimed Abilities(2 posts)→
More from Research
- Triangle Splatting SLAM: Imperial College's ECCV 2026 dense RGB-D SLAM with on-the-fly mesh extraction — rsasaki0109 · 2026-09-30
- Manifold opens early access: robotics eval platform runs thousands of GPU-parallel rollouts in 30 mins — paigeinsf · 2026-09-30
- 1,000 AI agents discover new CRISPR-like system in virus DNA within 24 hours — CurieuxExplorer · 2026-09-30
- Explaining just 5% of token positions retains nearly all audit success across 4.7M explanations — aisilab · 2026-09-30
- NTU's Persistence Forcing hits FID 1.63 on ImageNet 256 by heterogeneous refinement in pixel-space DiTs — NanyangTechnologicalUniversity · 2026-09-30
- IBM's Q&D trains proactive agents to ask better questions, beating a 15x larger model — ibm · 2026-09-30