One-Shot OPD: a single query keeps on-policy distillation improving for hundreds of steps

ParshinShojaee · x · 2026-09-04

Following their earlier Rethinking OPD work, the team dug into why On-Policy Distillation works and landed on One-Shot OPD: training on a single query keeps the student improving for hundreds of steps and recovers most of the gains of full-data OPD.

They frame OPD as "data-overfed but algorithm-starved":

Original post →

More from Research

Research channel →