Post-Training Benchmarks Are a Mess: Why Centralized Eval Orgs Distort Model Development

SonglinYang4 · x · 2026-09-05

A thread-length take on post-training benchmarks: every benchmark saturates, and LLM eval has always been ad hoc—settings were gamed via temperature/topp/length, now via harness/setting. Any decentralized eval is cherry-pickable noise, so centralized orgs like AA became the de facto standard, forcing models to optimize for leaderboard optics over real feel and working quality.

The author splits job tasks into three buckets: already AI-trainable, un-trainable by AI, and currently being trained (all Computer Use tasks). Crowdsourced frontier benchmarks (OSWorld v2, ALE, TB4/TB-Sci) form the closure of the third bucket and will steer base-model optimization for months—but they're the noisiest, and without accountability for error rates, "broken benchmarks" like GPQA-Diamond and HLE will keep multiplying.

Original post →

More from Models

Models channel →