31,352 repeated benchmark runs show LLM scores drift 3x more across days than within a day

ionutvi · reddit · 2026-09-07

A benchmark platform founder argues LLM evaluation should be treated as a longitudinal measurement problem, not a leaderboard problem, since API-served models change behind the scenes without version transitions.

Key findings:

Methodology:

The author invites criticism on time-series units, distinguishing drift from infrastructure effects, benchmark secrecy boundaries, and alternatives to change-point detectors.

Related event: 31,352 Repeat Tests Show LLM Benchmark Scores Drift Heavily Day to Day(2 posts)→

Original post →

More from Models

Models channel →