On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar
cs.LG
2026-09-28
Isolating rollout policy: no on-policy edge in accuracy, forgetting or sparsity; KL direction sets performance, learning rate sets forgetting; only harder-task transfer gains.
On-policy data has been credited with three virtues in post-training: less catastrophic forgetting, sparser parameter updates, better generalisation. The evidence mostly comes from comparing SFT with RLVR, and the two differ in objective, reward signal, supervision density and optimisation all at once, so the effect of rollout policy (which model generates the training trajectories) has never been isolated. The cost gap is practical: on-policy training generates fresh trajectories throughout training, while off-policy reuses a fixed dataset. A Cambridge group builds the controlled testbed the debate needs: strong-to-weak distillation with everything except the rollout policy held fixed.
A Llama-3.1-8B teacher is distilled into a Llama-3.2-1B student (Qwen2.5-7B into 1.5B as a control), across medical MedReason, scientific Science and arithmetic Countdown, 150 steps of full fine-tuning, three seeds. The loss is the token-level KL between student and teacher along rollout trajectories.
The key move breaks a default coupling. Off-policy distillation (OffPD) conventionally pairs with forward KL and on-policy distillation (OnPD) with reverse KL; the pairing is an accident of the sequence-level KL chain rule and single-sample estimation, not an optimality guarantee. Computing the KL over the full vocabulary makes rollout policy and KL direction independent axes, crossed with learning rates of 1e-5 and 5e-5, plus a sweep up to 6e-5.
Gradient analysis explains the asymmetry. The forward-KL gradient with respect to student logits is πS(v) − πT(v), bounded per coordinate in [−1, 1], so any change induced by moving the rollout policy is bounded linearly by the total-variation distance between trajectory distributions. The reverse-KL gradient is weighted by πS(v) times an unbounded log-ratio: it blows up when the student puts mass on tokens the teacher finds nearly impossible, and it has no gradient to recover teacher modes the student omits. Reverse KL therefore depends on training along the student's own prefixes. A rollout spectrum πλ fills in the space between, λ = −1 recovering OffPD and λ = +1 OnPD, with |λ| > 1 extrapolating toward tokens one model favours.
| Comparison | Metric | Result |
| OnPD vs OffPD | Best mean ID accuracy, 3 tasks | 72% vs 73% |
| Forward vs reverse KL | ID accuracy range | 71–73% vs 35–72% |
| LR 1e-5 vs 5e-5 | OOD change | within ±1.3 pts vs −11.2 to −14.0 pts |
| LR 1e-5 vs 5e-5 | Update sparsity (τ=10⁻⁶) | 85.3–89.6% vs 51.9–60.0% |
| λ>0 vs off-policy | Countdown-4E pass@k | +10–15 pts, consistent |
Rollout policy itself barely registers. KL direction sets task performance: forward KL sits at 71–73% across every rollout policy and learning rate, while reverse KL swings between 35% and 72% and is highly learning-rate sensitive. Learning rate sets forgetting and sparsity: update sparsity falls roughly linearly as learning rate grows (R² of 0.94 and 1.00) and OOD accuracy declines with it; at matched learning rates OffPD forgets less in nearly every comparison, and its updates are at least as sparse in every pairing. The claim that on-policy data yields sparser updates does not survive this.
The spectrum experiment matches the gradient analysis: forward KL stays above 80% accuracy across the whole spectrum, moving only 5.2 points at learning rate 1e-5, while reverse KL varies widely, works only on student-favoured rollouts and occasionally collapses at the higher rate.
On-policy does help in one place. On the harder Countdown-4E variant, more on-policy rollouts give 10–15 points higher pass@k than off-policy, under both KL directions. That advantage does not survive subsequent RLVR: reverse-KL checkpoints climb fast and then collapse, and the highest sustained reward on Countdown-4 comes from off-policy checkpoints (forward KL at the low rate, or reverse KL at the high rate) despite starting near 0% accuracy. An incidental finding: with a teacher instructed to reason in Spanish, only OnPD with reverse KL keeps the student answering in English; the other configurations transfer almost entirely.
Three robustness checks hold up: sampled KL estimators, no gradient clipping (reverse KL deteriorates sharply without it, forward KL holds), and Numina–MATH with 622-token average responses. On the long-reasoning set OnPD with reverse KL posts the best MATH-500 score, 59.0% against a 53.1% baseline, but that condition has one run, which the authors flag as suggestive.
The actionable part: on-policy distillation work should include OffPD as a standard baseline. It needs no live generation during training, and the savings are real compute. At the recipe level: to preserve old capabilities, forward KL with a small learning rate reaches top accuracy while keeping forgetting within 1.3 points; for output diversity, forward KL gives clearly higher pass@10 than reverse KL at matched pass@1; to keep the student from absorbing the teacher's style, OnPD with reverse KL is the one configuration that does it. Most of what got attributed to on-policy data is KL direction and learning rate doing the work.
Author-stated bounds: students at most 1.5B parameters, reasoning traces at most 2,000 tokens, a frozen teacher, and only same-family pairs. 150 training steps is short; long schedules are untested. Checkpoint selection on MATH-500 reuses the reported evaluations, which the authors admit may be optimistic. The reward collapse after RLVR is observed but not explained, and the 300-step RLVR window is short. The controlled design also bounds the conclusions: it isolates rollout policy within distillation, while the SFT-versus-RLVR debate also involves reward signals and supervision density, untouched here.