u-OPSD: Self-Distillation Without Labels or Teachers Beats GRPO on Math Reasoning

burny_tech · x · 2026-08-12

u-OPSD is a novel self-distillation method for language models. It requires no external labels, verifiers, or teacher models. Instead, it enables a model to improve itself by majority-voting across its own multiple rollouts and then distilling knowledge specifically on the disagreements.

Experiments show that this approach outperforms both supervised OPSD and GRPO on mathematical reasoning tasks.

Original post →

More from Research

Research channel →