On-policy distillation explained: teachers grade every token of student trajectories
helloiamleonie · x · 2026-09-25
Leonie walks through on-policy distillation (OPD): the student's own trajectory is passed to a teacher model, which grades each token by its own probabilities — measuring how 'surprised' it is at each step. Compared to sparse verifiable rewards, this gives the student a much denser, richer learning signal.
Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→
More from Research
- Nature paper: psilocybin reshapes latent temporal structure in brain activity — adeelrazi · 2026-09-25
- Francis Bach Blog: Taming Exploding Variance of Exponential Means With Least Squares — BachFrancis · 2026-09-25
- William & Mary AI Frontier Lab lands 5 NeurIPS acceptances — jindong_wang92 · 2026-09-25
- RUC open-sources EvoOntology, a self-evolving ontology layer for data agents via MCP — JeremyCMorgan · 2026-09-25
- Inside LFM2.5-2.6B's post-training recipe: SFT, RL, multi-domain distillation — helloiamleonie · 2026-09-25
- Transferring Qwen3.8's n-gram memory into a 0.8B model cuts perplexity 5.05% — Nicolodeva · 2026-09-25