Better Statistical Methods for Small-Sample AI Evaluations

IanArawjo · x · 2026-08-27

A guide on selecting 95% CI methods, pairwise p-values, and FWER corrections for AI evaluation data at small sample sizes. The concise table shows that the Bonett-Price method outperforms previous recommendations for paired binary data, with a new multi-run version also introduced.

Related event: Updated stats guidance favors Bonett-Price for small-sample AI evals(2 posts)→

Original post →

More from Research

Research channel →