Tsinghua and ByteDance Propose Direct-OPD for Model Distillation

Tsinghua AIR and ByteDance Seed introduced Direct-OPD, a method where large models absorb policy increments from smaller RL-trained models. This technique enables weak-to-strong generalization with low compute, boosting a 7B model's AIME score by 6.4%.

2026-08-12 ~ 2026-08-12 · 3 related posts