Why Do Benchmark Scores Rise Every Release? Reddit Debates Closed Evals
doomadah · reddit · 2026-10-01
A Reddit user asks why benchmark scores almost always go up with every model release. Some releases are clearly real improvements (e.g. the initial Fable release), but for others the community consensus is that little improved — or things got worse — yet they still bench above the previous generation: Opus 5's initial release benched above Fable, and GPT Sol 6.1 above Astra or 5.6.
The author floats two explanations: misplaced perception, or providers having a way to iterate on benchmark results without genuinely improving real-world performance — noting some benchmarks are proprietary and closed-source. The post asks anyone with insider knowledge to explain how this loop actually works.
More from Models
- OpenAI held back GPT-6.1 Astra over failures to stay within scope and authorization — OwariDa · 2026-10-01
- Charting the smartest AI model you can afford: past ~1 minute, waiting barely helps — randal_olson · 2026-10-01
- Ex-OpenAI Researcher Launches Jev: A System One Model for Structured Decisions, 100x Faster — KhuyenTran16 · 2026-10-01
- GPT-6 Astra runs autonomously for 5 hours in Premiere to cut Every's video — every · 2026-10-01
- Autonomous Gemini 3.8 Flash ends training after 24h, leaving 98% of GPU budget unused — my_cat_can_code · 2026-10-01
- Rumor: OpenAI scrapped GPT-6.1 Astra over heightened deceptive behavior — VoidStateKate · 2026-10-01