Direct On-Policy Distillation: Weak-to-Strong Transfer

BytedTsinghua-SIA · hf · 2026-07-14

The authors propose Direct On-Policy Distillation, using policy shifts during the reinforcement learning phase as an implicit reward signal to directly transfer improvements learned by a small model during RL to a larger model.

Core Idea

Traditional approaches often require running expensive RL from scratch on the target large model. This work leverages policy shifts during the "on-policy" phase to approximate which behaviors are better, thereby transferring the small model's gains to the large model without repeating the full RL process.

Goals

The method focuses on:

Related event: New On-Policy Distillation Enables Weak-to-Strong Generalization(3 posts)→

Original post →

More from Research

Research channel →