Study Finds 1 in 5 AI Benchmarks Regress in Newer Models
gleech · x · 2026-08-27
A new analysis by Paradigm 3 reveals that frontier AI models "sometimes get worse." Across 5,000+ model-benchmark pairs, about a fifth of benchmark scores declined in successive model iterations, with a median drop of 6.5% of the benchmark's range.
Key Findings
- Non-uniform regression: Regressions are rarer in coding and STEM tasks but more common in writing and creativity.
- Survivorship bias: More prominent or economically valuable benchmarks are less likely to regress. This likely reflects a selection filter where model checkpoints that regress on key metrics are shelved.
- Writing quality case: Analysis of Claude Opus models suggests Opus 4.8 is generally worse than predecessors on proxy measures, and Opus 5 shows no improvement on human-evaluated benchmarks.
- Explanations ruled out: The study rules out evaluation artifacts (QRPs), alignment tax damage, or strategic sandbagging as primary causes.
Related event: One in Five AI Benchmarks Regresses in Newer Models(2 posts)→
More from Models
- Mollick Warns Against Anthropomorphizing Agents in METR's HF Report — emollick · 2026-08-27
- Zai open-sources GLM-5.3-Flash: 320B-parameter model running on Chinese chips — ccerrato147 · 2026-08-27
- Perplexity Computer Launches: Local/Cloud Support, Claims 85.4% Benchmark — ChrisUniverse · 2026-08-27
- Critique of Claude style: Verbose but information-dense — TheZvi · 2026-08-27
- GLM-5.3 Flash review: Fixes 13 bugs with high cost-performance — PawelHuryn · 2026-08-27
- Opus 5 writing regression confirmed by benchmarks — gleech · 2026-08-27