Revamping AI Evaluations with IRT
AI Engineer · youtube · 2026-07-12
This talk argues that AI evaluations should move beyond the traditional "number of correct answers" approach and instead adopt Item Response Theory (IRT) from psychometrics.
Key takeaways include:
- IRT places each question on a unified scale with a margin of error, distinguishing effective questions from noise.
- Adaptive testing can measure the same capabilities with fewer questions, enabling the creation of rotatable, private benchmarks that are harder to contaminate.
- Question-fit statistics can expose data leakage/answer exposure, offering deeper insights into what a model has actually learned compared to a single score.
Speaker Alejandro Vidal believes these methods will make evaluations cheaper, more resistant to gaming, and better at revealing a model's true capability boundaries and knowledge gaps.
Related event: Revamping AI Evaluation with Item Response Theory(2 posts)→
More from Research
- VidMap uses RoMa coarse matching on all frames, fine-scale only for keyframes — ducha_aiki · 2026-09-11
- Bug Hunt Bench author: leaderboard noise is about 2-3 points — PawelHuryn · 2026-09-11
- Bug Hunt Bench ranks frontier coding models on 105 planted real-repo bugs — PawelHuryn · 2026-09-11
- PNAS paper shows a tiny billiard-ball system is a universal computer — undecidability lives in two dimensions — eigensteve · 2026-09-11
- New paper: Absolute pose estimation from affine cues and gravity direction — ducha_aiki · 2026-09-11
- LoMa Paper Ships REALLY HardPairs Dataset, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11