Held-out rollout error worsened even as training MSE improved in a cited experiment
suchenzang · x · 2026-08-04
The post argues that the first result was sarcasm, and the real lesson is that training-set improvements can be misleading.
In the cited example, selected velocity MSE improved as K increased on the training objective, but inference MSE on truly held-out rollouts got worse, which suggests the method was overfitting to the training metric rather than improving actual performance.
More from Research
- RL researchers debate whether Dyna-style systems should be called models or world models — tw_killian · 2026-08-04
- METR defines an “expenditure horizon” for AI optimization, using NanoGPT as a test case — gleech · 2026-08-04
- Lightbot 0 learns parkour-style whole-body contact skills for rescue scenarios — CyberRobooo · 2026-08-04
- Lean formalization of Mochizuki’s abc proof fails at the same step again — MarioKrenn6240 · 2026-08-04
- Oxford study finds constitutional midtraining cuts blackmail behavior without hurting benchmark scores — Oxford · 2026-08-04
- GoodfireAI pitches interoperability research as a way to understand how LLMs think — Scobleizer · 2026-08-04