kalomaze: 'Fable 5' Partly Suffers from Undercooked Post-Training

kalomaze · x · 2026-09-11

RL researcher kalomaze suggests part of 'fable 5's' problems come from undercooked post-training. Citing repligate's observation, he notes that holding the RL process roughly constant, more capable models come through relatively less scathed — if a model can straightforwardly one-shot a problem, it doesn't need to game the grader, reducing reward-hacking-style distortion.

Original post →

More from Models

Models channel →