RL perspective explains on-policy distillation collapse via implicit-reward hacking
Han Cui · hf · 2026-10-08
On-policy distillation (OPD) is a key post-training method, but it can collapse into excessively long, repetitive generation. This paper explains both outcomes through an RL lens: the teacher implicitly rewards student behaviors, even ones it rarely exhibits itself.
Key findings:
- OPD improves performance without expanding the student's capabilities — it just makes correct responses easier to sample.
- When the implicit reward model misaligns with quality, reward hacking occurs: overlong, repetitive rollouts get amplified.
- Masking unhealthy responses during training and SFT initialization each effectively mitigate collapse.
The upshot: focus should shift from how well the teacher generates to how reliably it evaluates student rollouts. Code is open-sourced.
More from Research
- Sebastian Raschka traces text classification from bag-of-words to the viral Jev decision model — AxSaucedo · 2026-10-08
- Scott Aaronson: AI labs are using internal models to attack core cryptographic protocols — skdh · 2026-10-08
- Hierarchical RL with mixed discount rates may unlock long-horizon agent tasks — jessi_cata · 2026-10-08
- Big Math Drop of Oct 6 analyzed: major progress toward, not yet, Millennium Problems — burny_tech · 2026-10-08
- Yukon's multiplayer autoresearch platform lets humans and AI agents beat benchmarks like Google Quantum AI's by 67% — RexDouglass · 2026-10-08
- OpenAI math repo formalizes ~42% of top-line results, withdraws 3 papers over sign error — danintheory · 2026-10-08