evalstats: Open-Source Tool for LLM Judge Bias-Corrected Stats Tests
IanArawjo · x · 2026-07-30
Developer Ian Arawjo announced an update to the evalstats package, introducing bias-corrected statistical tests for LLM judges.
- Core Features: Provides rigorous statistical analysis for LLM evaluations, from model and prompt comparisons to statistical inference.
- Bias Resilience: Specifically designed for small sample data regimes, resilient to biases introduced when using LLMs as judges.
- Availability: The defaults are battle-tested in Monte Carlo simulations, and researchers can use it today to analyze mixed human-AI subjects study data.
More from coding & agent
- LlamaIndex Founder: Humans May Stop Reviewing AI Code in 1-2 Years — dotey · 2026-07-30
- The New MCP Is Stateless: A Visual Guide to the Protocol Rewrite — dima806_dima · 2026-07-30
- ARC-AGI-3 API clarification: 'reasoning' field logs model output, not private CoTs — GregKamradt · 2026-07-30
- Claude Opus 5 Recreates Fallout Game in a Single 2.3MB HTML File — chrisfirst · 2026-07-30
- CodePilot v0.62 Adds Right-Click File Tree and Live Markdown Preview — op7418 · 2026-07-30
- A Practical Guide to Building Frontier-Lab Quality AI Evaluations — aakashgupta · 2026-07-30