evalstats Tool Offers Paper-Ready Model Eval Plots with Statistical Significance
IanArawjo · x · 2026-08-12
A developer shared the .plot() method of the open-source evaluation tool evalstats. Designed for paper-ready figures, it defaults to significance tiers rather than gradient colors: it highlights models statistically dominated by alternatives (at alpha=0.05) in red and the 'unbeaten' tier in blue. All confidence intervals are FWER corrected, making it ideal for comparing model performance on datasets like AI4Math.
More from Research
- OpenAI Models Reason in 'Alien Language', Making CoT Monitoring Nearly Impossible — basedjensen · 2026-08-12
- Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution — LMU · 2026-08-12
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design — Qing Zong · 2026-08-12
- SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents — Xiaofan Bai · 2026-08-12
- DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows? — Mizanur Rahman · 2026-08-12
- Latent-to-4D: Generating Reusable 4D Worlds Directly from Video Latents — Zihao Liu · 2026-08-12