31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day
ionutvi · reddit · 2026-09-07
A benchmark platform founder argues LLM evaluation should be treated as a longitudinal measurement problem, not a leaderboard problem, since API-served models change behind the scenes without version transitions.
Key findings:
- 31,352 repeated score observations across 49 models
- Within-day score std dev: 2.80 points; between-day daily median std dev: 8.43 points — roughly 3:1
- The author stops short of claiming providers change models daily (too many confounders), but argues temporal variation must be measured, not treated as noise
Methodology:
- Versioned benchmark configs; only compare compatible longitudinal observations
- Repeated execution-based evals instead of LLM judges
- Separate availability failures from valid outcomes, track serving/version metadata, run change-point detection
- Addresses benchmark contamination by publishing methodology while withholding the live task bank
The author invites criticism on time-series units, distinguishing drift from infrastructure effects, benchmark secrecy boundaries, and alternatives to change-point detectors.
Related event: 31,352 Repeat Tests Show LLM Benchmark Scores Drift Heavily Day to Day(2 posts)→
More from Models
- Astra's AGI estimate jumps with tool use — is 'ASI already here' just a harness question? — kevinnbass · 2026-09-07
- GLM 5.3 and Qwen 3.8 now run really well locally on single desktops — jasonkneen · 2026-09-07
- New benchmark probes LLM self-modeling: RL lifts open models but counterfactual errors persist — dair_ai · 2026-09-07
- Alexandr Wang flags Muse Spark 1.3 eval: time horizon now matches GPT-5.6 Sol and Opus 5 — alexandr_wang · 2026-09-07
- Cool presentation aside, Astra still can't nail research-level single-step reasoning — xiaosun86 · 2026-09-07
- GPT-6 Astra generates missile evasion simulation, showcasing stunning capability — algo_diver · 2026-09-07