evalstats Project Finds Most CI Methods Are Unreliable for Small-Sample AI and HCI Evaluations
IanArawjo · x · 2026-09-04
Researcher Ian Arawjo's evalstats project began as an effort to identify which confidence-interval methods work for small-sample AI evaluations, and revealed that HCI research — which routinely operates on small samples — needs these recommendations most. A follow-up study on pairwise CIs for between-subjects data is underway.
More from Research
- Mol-JEPA: A Multimodal JEPA Foundation Model for Molecules, One Year in the Making — TerribleAntelope9348 · 2026-09-04
- Two Years On: Six Guidelines for Making Research Impact via Open-Source in AI — lateinteraction · 2026-09-04
- GPT-6 Astra claims SOTA on ARC-AGI-3 at 66%, up from Sol's 8% — teortaxesTex · 2026-09-04
- Steering Qwen along a grader-vs-human dimension oddly shifts its personality — voooooogel · 2026-09-04
- CMU's AI Reviewer Beats Best Human Reviewer, Featured by Science — AkariAsai · 2026-09-04
- Chollet: ARC-AGI-4 lands Q1 2027, and solving ARC-3 is not AGI — fchollet · 2026-09-04