Adapting validity testing to benchmarks with IRT-based item-level analysis
sanmikoyejo · x · 2026-09-29
To adapt convergent and discriminant validity to benchmarks, the team tested whether benchmarks purporting to measure similar concepts correlate and dissimilar ones do not, asking the same questions at the item level using IRT models for finer-grained validity checks.
Related event: Psychometric Validity Framework Proposed for Evaluating AI Benchmarks(2 posts)→
More from Research
- Researchers argue catastrophic forgetting drives why fine-tuned bad behaviors persist in stronger models — QuintinPope5 · 2026-09-29
- HalluWorld Benchmark, Accepted at NeurIPS, Shows Models Nail Perception but Fail Simulation — xennygrimmato_ · 2026-09-29
- Quintin Pope: backdoor-style training setups are a poor stand-in for hypothesized inner optimizers — QuintinPope5 · 2026-09-29
- Ocular microtremors at ~100Hz may give the visual cortex apparent super-resolution — docmilanfar · 2026-09-29
- AI-designed viruses from scratch: 302 phage genomes synthesized, 16 worked — CallRevolutionary894 · 2026-09-29
- Reddit thread: softmax has only N-1 degrees of freedom — drop one input? — Kinexity · 2026-09-29