Guide to Statistical Methods for Small-Sample AI Evals
IanArawjo · x · 2026-08-27
This post provides recommendations for 95% CI methods, pairwise p-values, and FWER corrections for AI evals at small sample sizes. It identifies Bonett-Price as outperforming previous recommendations for paired binary data and introduces a multi-run version.
Related event: Updated stats guidance favors Bonett-Price for small-sample AI evals(2 posts)→
More from Research
- Flow matching decouples generative modeling from noise process — CSProfKGD · 2026-08-27
- Study finds Agent skill injection may lower Pass@2 rates — rohanpaul_ai · 2026-08-27
- Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig) — 1ncehost · 2026-08-27
- METR's new eval report gains traction over models losing track of tasks — isidentical · 2026-08-27
- Why Is the P=NP Question So Relevant in the AI Era? — yoavgo · 2026-08-27
- Fine-Tuning Guide: How Mistral 7B Saved $300k Over Foundation Models — Nice-Dragonfly-4823 · 2026-08-27