Paper: One training query recovers 72% of full-data on-policy distillation gains
heghbalz · x · 2026-09-06
A new paper, Rethinking On-Policy Distillation of LLMs II: One Training Example, probes the data-minimal limit of on-policy distillation (OPD):
- One query is enough to keep improving: one-shot OPD trains for hundreds of steps and recovers most of the full-data OPD gain across task domains and model families.
- State coverage explains it: a single query's rollouts already reach 71.5% of the states visited by full-data OPD, mostly within the first 100 steps; 16 semantically diverse queries hit 98.9% coverage and match full training.
- The bottleneck is algorithmic, not data: the student absorbs the teacher's dense token-level supervision slowly regardless of data volume. OPD is "data-overfed but algorithm-starved."
- Findings extend to multi-teacher OPD; content-light templates and off-domain queries also approach the real-query baseline.
Authors include Zhiyuan Liu and Ning Ding. Relevant to anyone paying for distillation datasets.
Related event: One training example recovers 72% of on-policy distillation gains(2 posts)→
More from Research
- AI editing a math paper spots a counterexample to a proposition the author planned to cite — jessi_cata · 2026-09-06
- Limits to narrow LLM complementarity: why 'taste is the bottleneck' won't hold — zetalyrae · 2026-09-06
- NeurIPS 2026 Registration Opens: Sydney Main Site Plus Atlanta and Paris Satellites — NeurIPSConf · 2026-09-06
- PolyU-Led Survey Maps Human-Centric AI: 3 Perspectives, 6 Layers from Body Perception to Embodied Agents — 机器之心 · 2026-09-06
- Terence Tao Clarifies Navier-Stokes Rumor Was Misread Hypothetical; Clay Still Lists Problem Unsolved — 机器之心 · 2026-09-06
- Interp vs mech interp: you can explain models without hunting for mechanisms — voooooogel · 2026-09-06