OPSA boosts AIME24 by 35 points using self-entropy without teacher distillation

heghbalz · x · 2026-09-01

New research questions on-policy distillation, suggesting gains come from suppressing unlikely tokens rather than teacher knowledge. The authors introduce OPSA (On-Policy Self-Adaptation), a teacher-free method that assigns entropy-adaptive negative advantages to low-probability tokens. This self-supervised approach achieved a +35 point improvement on the AIME24 benchmark.

Original post →

More from Models

Models channel →