One-Shot OPD: a single query keeps on-policy distillation improving for hundreds of steps
ParshinShojaee · x · 2026-09-04
Following their earlier Rethinking OPD work, the team dug into why On-Policy Distillation works and landed on One-Shot OPD: training on a single query keeps the student improving for hundreds of steps and recovers most of the gains of full-data OPD.
They frame OPD as "data-overfed but algorithm-starved":
- Data: one query already covers most of the states full data would visit.
- Algorithm: the student absorbs less of the remaining gap with each step, so returns diminish.
More from Research
- Researcher teases dynamic composite eval index as "evals run on Twitter vibes" — evijit · 2026-09-04
- Turning agent traces into training data: capture, sampling, and labelling, worked through — spilldahill · 2026-09-04
- Researcher Visualizes Qwen 2.5 Embedding Space, Says Mainstream Models' Conversational Ability Is Flattening — arianaram · 2026-09-04
- DeepMind's Prateek Jain on MatFormer: agents that dial compute up or down by task difficulty — jainprateek_ · 2026-09-04
- Fei-Fei Li on World Labs' Atlas: new view prediction as a world model primitive — a16z Podcast · 2026-09-04
- Embodied AI dataset ACE-Data-0 hits HF trending with ~30K downloads one week after release — liuziwei7 · 2026-09-04