Benchmark score reports need 95% CI error bars, argues ML practitioner

rmcwhorter99 · x · 2026-10-01

Commenter rmcwhorter99 calls out AI labs for never including 95% confidence interval error bars in benchmark score reports: "everyone does this iirc — and adding them would raise my timeline's IQ by half a stdev."

The quip highlights a real methodological gap: model eval scores are typically reported without statistical uncertainty, so small score differences driven by sampling noise or prompt-set variance get treated as genuine capability gaps. A pointed take on eval methodology worth discussing.

Original post →

More from Models

Models channel →