Unsupervised On-Policy Self-Distillation Improves LLMs

UCSanDiego · hf · 2026-08-12

UC SanDiego published a new research paper introducing an unsupervised on-policy self-distillation method for large language models.

By leveraging internal consistency and majority-vote pseudo-solutions, this approach can correct confident errors without requiring any external supervision.

Original post →

More from Research

Research channel →