One training example recovers 72% of on-policy distillation gains

A paper shows that on-policy distillation with just one training example recovers 72% of full-dataset gains, suggesting OPD suffers from data surplus rather than data scarcity and that algorithms, not more data, are the bottleneck.

2026-09-04 ~ 2026-09-06 · 2 related posts