New On-Policy Distillation Enables Weak-to-Strong Generalization
Researchers introduced Direct On-Policy Distillation (TOP-D) to achieve weak-to-strong generalization, using RL policy changes as implicit rewards to stably transfer improvements from small to large models.
2026-07-13 ~ 2026-07-15 · 3 related posts
- TOP-D Stabilizes Policy Distillation — HKUSTGZ · 2026-07-13
- Direct On-Policy Distillation: Weak-to-Strong Transfer — BytedTsinghua-SIA · 2026-07-14
- On-Policy Distillation for Weak-to-Strong Generalization — _akhaliq · 2026-07-15