New paper: supervising just 1% of tokens can match full on-policy distillation, 0.1% sometimes suffices

jiank_uiuc · x · 2026-09-23

A new MBZUAI–Ant Group paper, "1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation," studies why useful teacher guidance in sparse on-policy distillation (OPD) can still yield noisy updates when gradients are estimated from a single sampled token.

Key points:

Paper and code are public.

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Models

Models channel →