Rethinking on-policy distillation: researchers propose OLIVE, letting students learn from teacher continuations
May_F1_ · x · 2026-09-27
With on-policy distillation work proliferating, researchers from UPenn, UW and MIT are rethinking what it's actually needed for and whether simpler alternatives could work better.
They're exploring OLIVE: let the student model try first, then learn from the teacher's continuation. The project is at an early exploratory stage, with no experimental results published yet.
More from Research
- Rufus-Air paper: ordering post-training by reward reliability builds competitive open LLM — burny_tech · 2026-09-27
- Researchers teach LLMs to find interesting theorems, boosting discovery 4.3x — burny_tech · 2026-09-27
- DeepMind's Economic Policy for AGI framework evaluates 11 interventions — CurieuxExplorer · 2026-09-27
- Contrastive World Models Learns World Models in Latent Space Without Pixel Prediction — burny_tech · 2026-09-27
- CoRL 2026 workshop on continually self-improving robots opens call for papers, due Sep 28 — PeterStone_TX · 2026-09-27
- Martin Casado recommends the best talk on in-context learning, a first-principles view of LLMs — AccBalanced · 2026-09-27