Sparse distillation paper: supervising just 0.1%-1% of tokens can match or beat full OPD

jiank_uiuc · x · 2026-09-23

A new arXiv paper, 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation, tackles a hidden inefficiency in sparse on-policy distillation (OPD): even useful teacher guidance produces noisy updates when the gradient is estimated from a single sampled next token.

Method

Results

Paper and code are available.

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Models

Models channel →