UIUC paper: distilling with just 0.1% of tokens can beat full on-policy distillation

jiank_uiuc · x · 2026-09-23

Researchers present 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation, introducing the Information-Efficiency Ratio (IER) to measure gradient-estimation reliability in on-policy distillation.

Key insight: the expected distillation gradient averages over all possible next tokens, but in practice it's estimated from a single sampled token — so updates stay noisy even when teacher guidance is useful. IER is derived from a signal-to-noise decomposition in information geometry (higher IER = lower relative gradient estimation error) and can be approximated with a small candidate set to rank tokens.

Results:

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Research

Research channel →