Reasoning Models Improve by Updating Low-Probability Tokens Without Teachers

burkov · x · 2026-09-02

Research finds that while teacher model feedback is often noisy, student models still improve. The benefit stems mainly from suppressing tokens the student deems unlikely. The proposed On-Policy Self-Adaptation (OPSA) method focuses updates on these low-probability tokens, achieving similar gains without a teacher.

Related event: Study: On-Policy Distillation Gains Come From Suppressing Low-Probability Tokens, Not Teachers(3 posts)→

Original post →

More from Research

Research channel →