Paper: One training query recovers 72% of full-data on-policy distillation gains

heghbalz · x · 2026-09-06

A new paper, Rethinking On-Policy Distillation of LLMs II: One Training Example, probes the data-minimal limit of on-policy distillation (OPD):

Authors include Zhiyuan Liu and Ning Ding. Relevant to anyone paying for distillation datasets.

Related event: One training example recovers 72% of on-policy distillation gains(2 posts)→

Original post →

More from Research

Research channel →