NVIDIA's LSPD brings RL tricks to policy distillation, cutting rollouts by 75%
nvidia · hf · 2026-09-29
NVIDIA researchers reinterpret on-policy distillation (OPD) through the lens of RL, linking the reverse-KL objective to KL-regularized policy optimization and introducing Least-Square Policy Distillation (LSPD), which imports optimistic exploration and off-policy data reuse from value-based RL.
Key points:
- The idealized formulation achieves a sharp O(log K) regret bound under online exploration;
- LSPD outperforms existing distillation baselines by an average of +1.59 points (Avg@16) across six math reasoning benchmarks and diverse teacher-student setups;
- Pass@k evaluations up to k=64 show it better preserves policy diversity;
- Its fully off-policy variant matches vanilla OPD using only the first 25% of rollout batches, greatly cutting sampling cost.
More from Research
- Tsinghua humanoid robot plays badminton with one policy from just 30 min of human motion data — ChongZzZhang · 2026-09-29
- Open-source xvr AI aligns live X-ray with 3D CT at submillimeter accuracy, published in Nature — Dr_Alex_Crimi · 2026-09-29
- Simulating Human Consciousness: Paper Maps a New Frontier for AI and Robotics — ugail · 2026-09-29
- Getting AI 'drunk' makes it more likely to break rules and spill secrets, UNSW study finds — gaganghotra_ · 2026-09-29
- 176.9B MoE squeezed to ~1.89 effective bpw: GSQ-RCO GGUFs run Coder build in 29.6GB — Loginhe · 2026-09-29
- GRPO with a judge model biases toward longer answers — maybe why LLMs write essays to simple prompts — djcows · 2026-09-29