31,352 repeated LLM measurements show day-to-day scores vary 3x more than within-day runs

ionutvi · reddit · 2026-09-07

The founder of AI Stupid Level shared a longitudinal analysis based on 31,352 repeated benchmark score observations across 49 models: within-day score standard deviation was 2.80 points, while between-day standard deviation of daily medians was 8.43 points — roughly 3x — suggesting a single benchmark score shouldn't be treated as a permanent property of a model.

The author stresses this is descriptive, not proof that providers silently change models: sampling noise, task composition, serving-side behavior, infrastructure, and benchmark changes can all produce drift. Their methodology includes:

A second core issue is benchmark contamination: a continuously running benchmark whose tasks are fully public gradually becomes part of the environment models are optimized against, motivating a split between methodological transparency and evaluation-set secrecy. The argument: LLM benchmarks should shift from leaderboards toward continuous observability systems. A methodology PDF is publicly available.

Related event: 31,352 Repeat Tests Show LLM Benchmark Scores Drift Heavily Day to Day(2 posts)→

Original post →

More from Models

Models channel →