Distillation study: KL direction and learning rate matter more than on-policy rollouts
CambUni · hf · 2026-10-03
A controlled strong-to-weak distillation study across Llama3 and Qwen2.5 families finds that rollout policy is not central: token-level KL direction shapes performance and coverage (forward KL robust to rollout policy; reverse KL favors student rollouts), while learning rate governs forgetting and update sparsity. On-policy data helps generalization on harder Countdown variants but the advantage doesn't reliably persist after RLVR — challenging the default preference for on-policy learning.
More from Research
- AC2: actor-critic with action chunking enables partial rollouts, trains faster than GRPO — ZeYanjie · 2026-10-03
- KAIST's World Observer gives world models movable panoramic eyes to track unseen regions — Scobleizer · 2026-10-03
- Ofir Press defends new bug-finding benchmark: training the behavior is fine, test-set contamination is not — OfirPress · 2026-10-03
- CAS Five-Year Plan gives AI for Science its own chapter, setting up a US-vs-China metascience bet — teortaxesTex · 2026-10-03
- Draft standard for autonomous labs emerges as AI-driven science heats up — Afinetheorem · 2026-10-03
- Papers with 'agent' in the title get cited more: AI lit search reads titles too — rajammanabrolu · 2026-10-03