Post-Training Benchmarks Are a Mess: Why Centralized Eval Orgs Distort Model Development
SonglinYang4 · x · 2026-09-05
A thread-length take on post-training benchmarks: every benchmark saturates, and LLM eval has always been ad hoc—settings were gamed via temperature/topp/length, now via harness/setting. Any decentralized eval is cherry-pickable noise, so centralized orgs like AA became the de facto standard, forcing models to optimize for leaderboard optics over real feel and working quality.
The author splits job tasks into three buckets: already AI-trainable, un-trainable by AI, and currently being trained (all Computer Use tasks). Crowdsourced frontier benchmarks (OSWorld v2, ALE, TB4/TB-Sci) form the closure of the third bucket and will steer base-model optimization for months—but they're the noisiest, and without accountability for error rates, "broken benchmarks" like GPQA-Diamond and HLE will keep multiplying.
More from Models
- Claimed GPT-6 Astra + H3 Max combo builds an insanely fast realtime Blender renderer — jfischoff · 2026-09-06
- Unverified leak: OpenAI's post-Astra model to ship as 'AGI' with real-time actions — imjustnewatai · 2026-09-06
- Early Astra hands-on: code quality and data analysis feel incremental, says developer — xeophon · 2026-09-06
- Providers wage price war over serving DeepSeek-v4-flash, user burns massive tokens for a few dollars — MaziyarPanahi · 2026-09-06
- Annoyed by 'maximum length' prompts? Refreshing the page lets you keep chatting — hi-sci-collab · 2026-09-06
- GPT-6 builds a full 3D Roguelike level in Godot using just 3% of a 20X quota — op7418 · 2026-09-06