TOP-D Stabilizes Policy Distillation

HKUSTGZ · hf · 2026-07-13

This work introduces Trust Region Policy Distillation (TOP-D), aiming to transform highly variable and unstable on-policy distillation into a stable training paradigm.

The key approach involves dynamically constructing a "proximal teacher" during training to theoretically control gradient variance. The authors provide a global convergence analysis and monotonic improvement bounds, demonstrating that the overall training process is more reliable and stable. In experiments, TOP-D significantly improved training stability, sample efficiency, and final performance on mathematical reasoning tasks. Importantly, it achieves this with almost no extra computational overhead, making it a strong alternative to OPD.

Related event: New On-Policy Distillation Enables Weak-to-Strong Generalization(3 posts)→

Original post →

More from Research

Research channel →