Benchmark score reports need 95% CI error bars, argues ML practitioner
rmcwhorter99 · x · 2026-10-01
Commenter rmcwhorter99 calls out AI labs for never including 95% confidence interval error bars in benchmark score reports: "everyone does this iirc — and adding them would raise my timeline's IQ by half a stdev."
The quip highlights a real methodological gap: model eval scores are typically reported without statistical uncertainty, so small score differences driven by sampling noise or prompt-set variance get treated as genuine capability gaps. A pointed take on eval methodology worth discussing.
More from Models
- Investor: Model companies pay benchmark firms to test multiple versions pre-launch — deedydas · 2026-10-01
- Researcher: Model knowing the copyright joke is the punchline 'is understanding' — technollama · 2026-10-01
- Researcher: Model's unprompted video-joke disclaimer is hard to explain without 'understanding' — technollama · 2026-10-01
- Talking math and physics with LLMs feels like working with an exam-acing savant — burny_tech · 2026-10-01
- One-line take: GPT-6.1 Sol is underrated, says AI commentator — mallow610 · 2026-10-01
- "Opus 5.5 is the cheapest model" — human hours saved beat token pricing — CamBrazy3 · 2026-10-01