Study Finds Confidence Interval Methods Unreliable for Small-Sample AI Evaluation
Ian Arawjo's evalstats project finds that confidence interval methods commonly used in small-sample evaluations—especially in HCI research—suffer from poor coverage and overconfidence, with bootstrap performing worst.
2026-09-04 ~ 2026-09-04 · 2 related posts
- evalstats Project Finds Most CI Methods Are Unreliable for Small-Sample AI and HCI Evaluations — IanArawjo · 2026-09-04
- evalstats: Bootstrap CIs Worst Offender, Under-Covering and Overconfident at Small Samples — IanArawjo · 2026-09-04