evalstats: Bootstrap CIs Worst Offender, Under-Covering and Overconfident at Small Samples
IanArawjo · x · 2026-09-04
Ian Arawjo shared the TLDR of his evalstats project: across the board, the CI methods HCI researchers typically reach for frequently under-cover and are overconfident at the small samples their studies use — with bootstrap CIs being the worst offenders.
More from Research
- SPACE cuts agent LLM calls by 78.9% while raising success rate on long-horizon tasks — dair_ai · 2026-09-04
- CoRL 2026 hosts first workshop on imperfect robotics data: failures, OOD, human surprises — RobobertoMM · 2026-09-04
- Baseten launches Base Labs, a research org for open-source AI with fully shared recipes — baseten · 2026-09-04
- Study: AI companions rival human friendship, users mourn forced separations — EricTopol · 2026-09-04
- MazeBench: New 3D Spatial Reasoning Benchmark Where Prior SOTA Agents Score Just 1% — patience_cave · 2026-09-04
- "World Models from Scratch": a hands-on open-source book launches its first release — Cohere_Labs · 2026-09-04