Why distillation gradients are noisy: single-token sampling breaks the expected-teacher estimate

jiank_uiuc · x · 2026-09-23

Technical follow-up from the same OPD paper author: the expected distillation gradient averages over possible next tokens, but in practice it's estimated from just one sampled token, making the update noisy even when the underlying teacher guidance is useful. The thread frames the problem in information geometry.

Related event: Sparse On-Policy Distillation: Supervising Just 1% of Tokens Can Match Full OPD(7 posts)→

Original post →

More from Research

Research channel →