Analysis of 31k LLM benchmarks: within-day variation 2.8 pts
ionutvi · reddit · 2026-08-29
Analysis of 31,352 hourly LLM benchmark scores reveals a within-day variation of 2.8 points versus an 8.4-point variation between days. The author open-sourced AIStupidLevel, a platform for continuous capability drift monitoring with 169,858 runs. It features four test suites and a smart router that switches models based on real-time performance metrics like stability and tool-calling reliability.
Related event: 31K Hourly Benchmarks Show LLM Scores Swing 3x More Across Days(2 posts)→
More from Infra
- Nvidia's Rubin Ultra reportedly downgraded to 8-High HBM4 — kevinsxu · 2026-08-29
- AI Latency Beyond the Model: Mapping 19 Full-Path Patterns — bibryam · 2026-08-29
- Jarvislabs Offers On-Demand H200 Clusters as GPU Access Gets Harder — algo_diver · 2026-08-29
- 31K hourly LLM benchmarks show 8.4-point day-to-day variation, 3x within-day noise — ionutvi · 2026-08-29
- Four practical ways to optimize end-to-end AI latency — bibryam · 2026-08-29
- Performance Optimization: Latency Reduced from 8ms to 0.87ms — DanielLockyer · 2026-08-29