One training example recovers 72% of on-policy distillation gains
A paper shows that on-policy distillation with just one training example recovers 72% of full-dataset gains, suggesting OPD suffers from data surplus rather than data scarcity and that algorithms, not more data, are the bottleneck.
2026-09-04 ~ 2026-09-06 · 2 related posts
- On-policy distillation improves for hundreds of steps from a single query — it's algorithm-starved, not data-starved — Thinking-Space · 2026-09-04
- Paper: One training query recovers 72% of full-data on-policy distillation gains — heghbalz · 2026-09-06