Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD

A late-September arXiv paper from MBZUAI, Ant Group and other institutions, "1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation" (arXiv:2609.24432), examines gradient estimation in sparse on-policy distillation (OPD). One of the authors, jiankuiuc, introduced the work in a series of posts: sparse OPD applies teacher supervision to only a small fraction of tokens in the student's generated trajectories, and experiments show that supervising just 1% of tokens matches full OPD — in some cases 0.1% suffices and can even surpass full supervision.

Confirmed

Why it matters

2026-09-22 ~ 2026-09-23 · 7 related posts

Primary sources

3 near-duplicate retellings: jiank_uiuc · jiank_uiuc · jiank_uiuc