On-policy distillation explained: teachers grade every token of student trajectories

helloiamleonie · x · 2026-09-25

Leonie walks through on-policy distillation (OPD): the student's own trajectory is passed to a teacher model, which grades each token by its own probabilities — measuring how 'surprised' it is at each step. Compared to sparse verifiable rewards, this gives the student a much denser, richer learning signal.

Related event: Liquid AI Releases LFM2.5-2.6B and Details Its On-Device Training Recipe(10 posts)→

Original post →

More from Research

Research channel →