New Self-Distillation Method Boosts LLM Self-Correction Without Supervision
burny_tech · x · 2026-08-08
Traditional self-distillation methods for LLMs often rely on gold answers, verifier rewards, or stronger teacher models. A new paper, "On-Policy Self-Distillation without Any Supervision," introduces a fully self-supervised approach.
The core mechanism involves:
- Sampling multiple solutions per problem
- Majority-voting to generate a pseudo-answer
- Distilling the pseudo-solution-conditioned distribution into the rollouts that disagreed
This provides dense on-policy corrections without labels. Experiments show that on Qwen3 math tasks in non-thinking mode, this method outperforms supervised OPSD and GRPO by 3.2 and 8.9 points, respectively.
More from Research
- PrimeIntellect Launches Multi-Agent Reinforcement Learning Training Stack — shi_weiyan · 2026-08-08
- FactorJEPA: A New World Model for Crowded Urban Environments — Kapil Wanaskar · 2026-08-08
- The Enduring Value of Data Hinges on the Future Cost of Verification and Generation — oyhsu · 2026-08-08
- HKU's Hengshuang Zhao Named MIT TR35 China for Work in Embodied AI and 3D Vision — YiMaTweets · 2026-08-08
- Genomic Intelligence to Demo DNA Models and Agent Integrations in Upcoming Webinar — julia_kiseleva · 2026-08-08
- AI-Generated Patches Fail Half the Time, Study of 6,000+ Patches Finds — WeldPond · 2026-08-08