31K hourly LLM benchmarks show 8.4-point day-to-day variation, 3x within-day noise

ionutvi · reddit · 2026-08-29

The author built a continuous evaluation pipeline and analyzed 31,352 hourly benchmark scores across 49 model identifiers (coding, deep reasoning, tool calling, high-frequency canaries):

The system is open-sourced as AIStupidLevel (MIT), totaling 169,858 benchmark runs and 88M+ tokens, monitoring 22 models across 6 providers, and powers an OpenAI-compatible router that selects models by current performance, stability, latency and cost.

Related event: 31K Hourly Benchmarks Show LLM Scores Swing 3x More Across Days(2 posts)→

Original post →

More from Infra

Infra channel →