31,352 repeated LLM measurements show day-to-day scores vary 3x more than within-day runs
ionutvi · reddit · 2026-09-07
The founder of AI Stupid Level shared a longitudinal analysis based on 31,352 repeated benchmark score observations across 49 models: within-day score standard deviation was 2.80 points, while between-day standard deviation of daily medians was 8.43 points — roughly 3x — suggesting a single benchmark score shouldn't be treated as a permanent property of a model.
The author stresses this is descriptive, not proof that providers silently change models: sampling noise, task composition, serving-side behavior, infrastructure, and benchmark changes can all produce drift. Their methodology includes:
- Repeated trials instead of single runs
- Execution-based scoring over LLM-as-judge
- Benchmark versioning: no silent comparison against old baselines after config changes
- Separating capability failures from provider/availability failures
- Recording systemfingerprint metadata, while noting an unchanged fingerprint proves nothing
- Cross-model correlation to attribute movement to provider-side changes
A second core issue is benchmark contamination: a continuously running benchmark whose tasks are fully public gradually becomes part of the environment models are optimized against, motivating a split between methodological transparency and evaluation-set secrecy. The argument: LLM benchmarks should shift from leaderboards toward continuous observability systems. A methodology PDF is publicly available.
Related event: 31,352 Repeat Tests Show LLM Benchmark Scores Drift Heavily Day to Day(2 posts)→
More from Models
- Mystery model Omen Alpha spotted; tokenizer tests point to new Zhipu GLM — realsohamparekh · 2026-09-07
- Training mixtures are now all synthetic: small-model training is really distillation — RexDouglass · 2026-09-07
- Sol High usage test: one complex prompt eats 5% of the 5-hour limit — remixedmoon5 · 2026-09-07
- Bodhan AI open-weights speech, vision and translation models for Indian languages on Hugging Face — selfawareatom · 2026-09-07
- Claude Max user reports a week of erratic usage-limit bugs and resets — tonimedic · 2026-09-07
- Dev Reminder: Astra Shines in Demo-Friendly Domains, but AGI Hinges on System-Level Understanding — Scobleizer · 2026-09-07