Developer open sources NoiseCheck, exposing statistical fallacies in model evals
Formal-King3851 · reddit · 2026-08-16
The author built a CLI tool called noisecheck to determine if model evaluation improvements are real or just statistical noise. Real-world tests on models like DeepSeek and GLM revealed that a "lead of 2.7 points" can be statistically meaningless, with rankings flipping upon scaling the sample size. Experiments also showed that a single model's score fluctuated by ±2.9 points across identical runs, exceeding the marginal gaps between different models.
The tool features cluster-aware resampling, judge consistency checks (GPT-4 vs. human experts Kappa 0.15), and power analysis, designed to act as a CI gate to prevent releases based on spurious data. The code is open source.
More from coding & agent
- BBOT Open Source Scanner Automates Bug Bounty Recon and ASM — tom_doerr · 2026-08-17
- CAKE Paper: Evolving Compiler Harness via AI for SOTA Kernels — lazowska · 2026-08-17
- GitHub Spec Kit: Define specs before letting AI agents build code — Shruti_0810 · 2026-08-17
- TradingAgents: Open-source multi-agent LLM trading framework — mdancho84 · 2026-08-17
- Making local models useful for coding: Hybrid cloud planning with local micro-patches — djpaul666 · 2026-08-17
- SpaceX Officially Closes Cursor Acquisition; Cursor Says It Now Has Access to 'Largest Fleet of GPUs in the World' — HaktanSuren · 2026-08-17