Auditing 162 LLM benchmark gaps: only 20 survive the noise check
maverick_man1111 · reddit · 2026-10-09
The author audited 162 "X beats Y" gaps from recent Claude/GPT/Gemini/Mistral launches and 9 leaderboards against each benchmark's sampling noise. Only 20 separate cleanly. Of 44 gaps quoted in six launch posts, 28 are unverifiable from published data (15 on Gemini 4 Argon, 9 on Mistral Large 4). On SWE-bench Multilingual, the 72.7% vs 66.3% lead is within the noise interval and six models are statistically indistinguishable from first place. Pipeline built with Claude Code; data and calculator are public.
More from Models
- Ai2: the next Olmo model is already training, fully open-model commitment unchanged — sewon__min · 2026-10-09
- Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages — solyarisoftware · 2026-10-09
- Nous Research's typo 'Hemres Agnet' looks like a Hermes Agent teaser — Teknium · 2026-10-09
- Study: Outdated Gemini 2.5 Advice Rated on Par With Doctors in Urgent Care, No Safety Issues — emollick · 2026-10-09
- Scott Aaronson: the AI skeptics' 'wake me up' warnings have all come true — scottleibrand · 2026-10-09
- Former multimodal researcher: aligning model 'taste' for everyone is unsolvable — A_K_Nain · 2026-10-09