Direct-OPD Transfers Weak Model RL Gains to Strong Models

tokenbender · x · 2026-08-21

The paper proposes Direct On-Policy Distillation (Direct-OPD) to avoid expensive RL runs on large models. It runs RL on a weaker model and transfers the policy shift as an implicit reward to the stronger student. Empirically, Direct-OPD consistently improves stronger target models, boosting Qwen3-1.7B accuracy from 48.3% to 58.3%.

Original post →

More from Models

Models channel →