Rethinking on-policy distillation: researchers propose OLIVE, letting students learn from teacher continuations

May_F1_ · x · 2026-09-27

With on-policy distillation work proliferating, researchers from UPenn, UW and MIT are rethinking what it's actually needed for and whether simpler alternatives could work better.

They're exploring OLIVE: let the student model try first, then learn from the teacher's continuation. The project is at an early exploratory stage, with no experimental results published yet.

Original post →

More from Research

Research channel →