Distillation study: KL direction and learning rate matter more than on-policy rollouts

CambUni · hf · 2026-10-03

A controlled strong-to-weak distillation study across Llama3 and Qwen2.5 families finds that rollout policy is not central: token-level KL direction shapes performance and coverage (forward KL robust to rollout policy; reverse KL favors student rollouts), while learning rate governs forgetting and update sparsity. On-policy data helps generalization on harder Countdown variants but the advantage doesn't reliably persist after RLVR — challenging the default preference for on-policy learning.

Original post →

More from Research

Research channel →