evalstats Project Finds Most CI Methods Are Unreliable for Small-Sample AI and HCI Evaluations

IanArawjo · x · 2026-09-04

Researcher Ian Arawjo's evalstats project began as an effort to identify which confidence-interval methods work for small-sample AI evaluations, and revealed that HCI research — which routinely operates on small samples — needs these recommendations most. A follow-up study on pairwise CIs for between-subjects data is underway.

Related event: Study Finds Confidence Interval Methods Unreliable for Small-Sample AI Evaluation(2 posts)→

Original post →

More from Research

Research channel →