Paper: Supervising Just 0.1%-1% of Tokens Can Match Full On-Policy Distillation

jiank_uiuc · x · 2026-09-23

The paper "1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation" shows that useful teacher guidance still yields noisy updates. The authors introduce the Information-Efficiency Ratio (IER) to measure gradient-estimation reliability and combine it with usefulness scores, matching or beating full OPD with only 0.1%-1% of tokens in their experiments. Paper and code are available.

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Research

Research channel →