Auditing 162 LLM benchmark gaps: only 20 survive the noise check

maverick_man1111 · reddit · 2026-10-09

The author audited 162 "X beats Y" gaps from recent Claude/GPT/Gemini/Mistral launches and 9 leaderboards against each benchmark's sampling noise. Only 20 separate cleanly. Of 44 gaps quoted in six launch posts, 28 are unverifiable from published data (15 on Gemini 4 Argon, 9 on Mistral Large 4). On SWE-bench Multilingual, the 72.7% vs 66.3% lead is within the noise interval and six models are statistically indistinguishable from first place. Pipeline built with Claude Code; data and calculator are public.

Original post →

More from Models

Models channel →