31K hourly LLM benchmarks show 8.4-point day-to-day variation, 3x within-day noise
ionutvi · reddit · 2026-08-29
The author built a continuous evaluation pipeline and analyzed 31,352 hourly benchmark scores across 49 model identifiers (coding, deep reasoning, tool calling, high-frequency canaries):
- Within-day variation was 2.8 points vs 8.4 between days (3×), suggesting isolated hourly moves are mostly stochastic noise while sustained daily-window changes are the stronger signal for performance drift.
- Coding answers are actually executed, not just model-judged; tool-calling is tested in isolated Docker environments; each task runs 5 times and is aggregated.
- The pipeline aggregates to daily medians and applies sequential change-point detection with minimum-effect thresholds before flagging degradation or recovery.
- At screenshot time it detected a 32% sustained decline in Gemini 3.1 Flash Lite, classified as critical.
The system is open-sourced as AIStupidLevel (MIT), totaling 169,858 benchmark runs and 88M+ tokens, monitoring 22 models across 6 providers, and powers an OpenAI-compatible router that selects models by current performance, stability, latency and cost.
Related event: 31K Hourly Benchmarks Show LLM Scores Swing 3x More Across Days(2 posts)→
More from Infra
- Tencent Hunyuan releases official 1-bit quantization for Hy4 with minimal accuracy loss — lakySK · 2026-08-29
- AI Master's Student Asks: Dual Tesla T10 vs Mi50 for Budget Rig? — fightingCookie0301 · 2026-08-29
- Google tightens Android memory rules amid AI-driven shortage — emmanuelvivier · 2026-08-29
- Nvidia's Rubin Ultra reportedly downgraded to 8-High HBM4 — kevinsxu · 2026-08-29
- AI Latency Beyond the Model: Mapping 19 Full-Path Patterns — bibryam · 2026-08-29
- Jarvislabs Offers On-Demand H200 Clusters as GPU Access Gets Harder — algo_diver · 2026-08-29