evalstats v0.2.6: Well-calibrated statistics for AI evals on small samples

IanArawjo · x · 2026-08-29

evalstats v0.2.6 is released. The updated compare() method produces well-calibrated statistics for AI evaluation results, even with small sample sizes or LLM-based judges (with limited human labels). A plot demonstrating the calibration is available in the linked release.

Original post →

More from coding & agent

coding & agent channel →