Study: On-policy distillation works by suppressing low-probability tokens
Purdue · hf · 2026-09-01
The paper "Does On-Policy Distillation Really Distill?" investigates the mechanism behind on-policy distillation. It finds that the process relies mainly on suppressing low-probability tokens rather than teacher guidance. This motivates a supervision-free entropy-adaptive method that substantially improves reasoning performance.
Related event: Study: On-Policy Distillation Works by Suppressing Low-Probability Tokens(2 posts)→
More from Research
- Claude 5.1 Rebuilds Venus Map from NASA Data, Boosting Resolution Significantly — haider1 · 2026-09-02
- Arena Launches 2026 Academic Partnerships Program with Up to $50k Funding per Project — arena · 2026-09-02
- UT PGE Faculty Discuss AI Strategy for Teaching and Research — GeostatsGuy · 2026-09-02
- Weekly Humanoid Papers: Sit/stand control waves and LAC research — carlosdponx · 2026-09-02
- Idea for alignment: agents should recognize impossible tasks — JacquesThibs · 2026-09-02
- Human Baseline on Mazebench: Top 10 Score Average 70% — patience_cave · 2026-09-02