Reasoning Models Improve by Updating Low-Probability Tokens Without Teachers
burkov · x · 2026-09-02
Research finds that while teacher model feedback is often noisy, student models still improve. The benefit stems mainly from suppressing tokens the student deems unlikely. The proposed On-Policy Self-Adaptation (OPSA) method focuses updates on these low-probability tokens, achieving similar gains without a teacher.
More from Research
- The Misconception that Almost Stopped AI: How Models Learn — burny_tech · 2026-09-02
- Why gradient descent works in high dimensions: The saddle point problem — burny_tech · 2026-09-02
- Analysis: Astra's recurrent depth could yield a 1.38x effective-parameter multiplier — scaling01 · 2026-09-02
- Proofs and Prompts: Mathematicians reflect on AI reshaping math — stevenstrogatz · 2026-09-02
- Fourier Neural Operators: Solving PDEs 1000x Faster — burny_tech · 2026-09-02
- Dallas Fed: Texas firms with higher GenAI exposure post 8-9% fewer jobs — rohanpaul_ai · 2026-09-02