New On-Policy Distillation Enables Weak-to-Strong Generalization

Researchers introduced Direct On-Policy Distillation (TOP-D) to achieve weak-to-strong generalization, using RL policy changes as implicit rewards to stably transfer improvements from small to large models.

2026-07-13 ~ 2026-07-15 · 3 related posts