Why Do Benchmark Scores Rise Every Release? Reddit Debates Closed Evals

doomadah · reddit · 2026-10-01

A Reddit user asks why benchmark scores almost always go up with every model release. Some releases are clearly real improvements (e.g. the initial Fable release), but for others the community consensus is that little improved — or things got worse — yet they still bench above the previous generation: Opus 5's initial release benched above Fable, and GPT Sol 6.1 above Astra or 5.6.

The author floats two explanations: misplaced perception, or providers having a way to iterate on benchmark results without genuinely improving real-world performance — noting some benchmarks are proprietary and closed-source. The post asks anyone with insider knowledge to explain how this loop actually works.

Related event: Reddit debates benchmark inflation and missing confidence intervals in AI model releases(2 posts)→

Original post →

More from Models

Models channel →