Tsinghua and ByteDance Propose Direct-OPD for Model Distillation
Tsinghua AIR and ByteDance Seed introduced Direct-OPD, a method where large models absorb policy increments from smaller RL-trained models. This technique enables weak-to-strong generalization with low compute, boosting a 7B model's AIME score by 6.4%.
2026-08-12 ~ 2026-08-12 · 3 related posts
- HF Journal Club: On-Policy Distillation Transfers RL Gains to Large Models at Low Compute — Hugging Face · 2026-08-12
- Direct-OPD: Transferring RL Policy Shifts from Small to Large Models — _lewtun · 2026-08-12
- Direct-OPD Paper: 7B Model Boosts AIME Score by 6.4% Using 1.5B Policy Shift — _lewtun · 2026-08-12