evalstats v0.2.6: Well-calibrated statistics for AI evals on small samples
IanArawjo · x · 2026-08-29
evalstats v0.2.6 is released. The updated compare() method produces well-calibrated statistics for AI evaluation results, even with small sample sizes or LLM-based judges (with limited human labels). A plot demonstrating the calibration is available in the linked release.
More from coding & agent
- GLM-5.3 Launches on Tinker with 256k Context — simonguozirui · 2026-08-29
- Conifer SDK Open-Sourced: Unified Gateway with Exact Cost Tracking — ycombinator · 2026-08-29
- Browser Use launches iMessage web agents for booking and shopping — _AustinCalvert_ · 2026-08-29
- AgentHeights Gamifies Agentic Orchestration with Virtual Office — edgarpavlovsky · 2026-08-29
- From Single Screen to Multi-Step Tasks: A 7-Step Roadmap for Medical AI Agents — MaryamMiradi · 2026-08-29
- Prime Agent: A Self-Improving RLM Harness for Coding and Autonomous Tasks — xeophon · 2026-08-29