31K Hourly Benchmarks Show LLM Scores Swing 3x More Across Days
An analysis of 31,352 hourly LLM benchmark scores across 49 model identities found that scores vary about 8.4 points across days—three times the within-day variation of 2.8 points. The author open-sourced the AIStupidLevel platform for continuous model monitoring.
2026-08-29 ~ 2026-08-29 · 2 related posts
- Analysis of 31k LLM benchmarks: within-day variation 2.8 pts — ionutvi · 2026-08-29
- 31K hourly LLM benchmarks show 8.4-point day-to-day variation, 3x within-day noise — ionutvi · 2026-08-29