NVIDIA's LSPD brings RL tricks to policy distillation, cutting rollouts by 75%

nvidia · hf · 2026-09-29

NVIDIA researchers reinterpret on-policy distillation (OPD) through the lens of RL, linking the reverse-KL objective to KL-regularized policy optimization and introducing Least-Square Policy Distillation (LSPD), which imports optimistic exploration and off-policy data reuse from value-based RL.

Key points:

Original post →

More from Research

Research channel →