On-policy distillation gains come from suppressing low-prob tokens, not the teacher

burny_tech · x · 2026-09-03

The paper "Does On-Policy Distillation Really Distill?" (Yi Ding, Ruqi Zhang) quantitatively dissects where on-policy distillation (OPD) gains actually come from, with surprising results:

The authors propose On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages: stronger signals at high-entropy positions, tail-token suppression, and even redistribution of probability mass among head tokens.

Results: versus base Qwen3-1.7B, OPSA improves Avg@32 on AIME24 by 35.41 points (263% relative gain), more than doubles Pass@32 across three benchmarks, and beats OPD by 16.77 points on AIME24 Avg@32, generalizing across model families and tasks.

Related event: On-policy distillation gains stem from suppressing unlikely tokens, study finds(4 posts)→

Original post →

More from Research

Research channel →