RL perspective explains on-policy distillation collapse via implicit-reward hacking

Han Cui · hf · 2026-10-08

On-policy distillation (OPD) is a key post-training method, but it can collapse into excessively long, repetitive generation. This paper explains both outcomes through an RL lens: the teacher implicitly rewards student behaviors, even ones it rarely exhibits itself.

Key findings:

The upshot: focus should shift from how well the teacher generates to how reliably it evaluates student rollouts. Code is open-sourced.

Original post →

More from Research

Research channel →