New Paper: Self-Distillation May Not Suit Thinking Models

amaarora · x · 2026-07-17

A new paper takes a closer look at on-policy self-distillation (OPSD) for thinking models.

The authors point out that OPSD was originally seen as a promising path for recursive self-improvement, as thinking models can leverage their own "privileged information" for reasoning, verification, and error correction. However, the experimental results were surprising: OPSD might actually degrade the performance of thinking models.

Original post →

More from Research

Research channel →