Better Statistical Methods for Small-Sample AI Evaluations
IanArawjo · x · 2026-08-27
A guide on selecting 95% CI methods, pairwise p-values, and FWER corrections for AI evaluation data at small sample sizes. The concise table shows that the Bonett-Price method outperforms previous recommendations for paired binary data, with a new multi-run version also introduced.
Related event: Updated stats guidance favors Bonett-Price for small-sample AI evals(2 posts)→
More from Research
- Conjecture: All formal artifacts to be vibe-formalized in Lean by 2026 — spikedoanz · 2026-08-27
- Flow matching decouples generative modeling from noise process — CSProfKGD · 2026-08-27
- Study finds Agent skill injection may lower Pass@2 rates — rohanpaul_ai · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27
- METR's new eval report gains traction over models losing track of tasks — isidentical · 2026-08-27
- Why Is the P=NP Question So Relevant in the AI Era? — yoavgo · 2026-08-27