Why distillation gradients are noisy: single-token sampling breaks the expected-teacher estimate
jiank_uiuc · x · 2026-09-23
Technical follow-up from the same OPD paper author: the expected distillation gradient averages over possible next tokens, but in practice it's estimated from just one sampled token, making the update noisy even when the underlying teacher guidance is useful. The thread frames the problem in information geometry.
More from Research
- Does sparse attention actually speed up video generation? Three conditions, measured — fruesome · 2026-09-23
- Basecamp Research raises $140m to design drugs from genetic data of 1m+ unstudied organisms — _rockt · 2026-09-23
- Basecamp Research Raises $140M From Nvidia and Anthropic's Anthology Fund to Turn Evolution Into Training Data — The Decoder · 2026-09-23
- Lean's type theory proves Con(ZF) from excluded middle alone — result found with AI — MikePFrank · 2026-09-23
- Mila's Blake Richards wins 2026 Schmidt Sciences Polymath Award for brain-and-algorithms research — Mila_Quebec · 2026-09-23
- When AI Writes DNA: Radical Numerics CEO on the Biosecurity Arms Race — Latent Space · 2026-09-23