OPD: Roles, Pathologies, and Regulation

Rui Wang · hf · 2026-07-17

The Roles, Pathologies, and Regulation of On-Policy Distillation

This paper systematically investigates the role and failure modes of On-Policy Distillation (OPD) in LLM post-training. The authors' core conclusion is that OPD acts more like an "exploration catalyst"—it pushes the student model toward correct reasoning paths via dense token-level guidance, but doesn't raise the capability ceiling.

Main Findings

Two Pathologies

Regulation Methods

The authors tested lightweight regulation methods, including:

These methods mitigate length speculation and make distillation more stable. Experiments covering 7 benchmarks show that regulated OPD consistently outperforms OPD variants and RLVR baselines.

Original post →

More from Research

Research channel →