Adapting validity testing to benchmarks with IRT-based item-level analysis

sanmikoyejo · x · 2026-09-29

To adapt convergent and discriminant validity to benchmarks, the team tested whether benchmarks purporting to measure similar concepts correlate and dissimilar ones do not, asking the same questions at the item level using IRT models for finer-grained validity checks.

Related event: Psychometric Validity Framework Proposed for Evaluating AI Benchmarks(2 posts)→

Original post →

More from Research

Research channel →