Guide to Statistical Methods for Small-Sample AI Evals

IanArawjo · x · 2026-08-27

This post provides recommendations for 95% CI methods, pairwise p-values, and FWER corrections for AI evals at small sample sizes. It identifies Bonett-Price as outperforming previous recommendations for paired binary data and introduces a multi-run version.

Related event: Updated stats guidance favors Bonett-Price for small-sample AI evals(2 posts)→

Original post →

More from Research

Research channel →