New Self-Distillation Method Boosts LLM Self-Correction Without Supervision

burny_tech · x · 2026-08-08

Traditional self-distillation methods for LLMs often rely on gold answers, verifier rewards, or stronger teacher models. A new paper, "On-Policy Self-Distillation without Any Supervision," introduces a fully self-supervised approach.

The core mechanism involves:

This provides dense on-policy corrections without labels. Experiments show that on Qwen3 math tasks in non-thinking mode, this method outperforms supervised OPSD and GRPO by 3.2 and 8.9 points, respectively.

Original post →

More from Research

Research channel →