Study: On-policy distillation works by suppressing low-probability tokens

Purdue · hf · 2026-09-01

The paper "Does On-Policy Distillation Really Distill?" investigates the mechanism behind on-policy distillation. It finds that the process relies mainly on suppressing low-probability tokens rather than teacher guidance. This motivates a supervision-free entropy-adaptive method that substantially improves reasoning performance.

Related event: Study: On-Policy Distillation Works by Suppressing Low-Probability Tokens(2 posts)→

Original post →

More from Research

Research channel →