1% of Tokens Can Be Enough: MBZUAI and Ant Group Improve On-Policy Distillation

jiank_uiuc · x · 2026-09-23

A MBZUAI–Ant Group paper (arXiv:2609.24432) studies gradient estimation in sparse on-policy distillation (OPD), where teacher supervision is allocated to a small subset of tokens:

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Research

Research channel →