On-Policy Self-Distillation: LLMs self-teach with 4-8x token efficiency over GRPO

burny_tech · x · 2026-09-13

Researchers from UCLA, HKU, and Meta Superintelligence Labs introduce On-Policy Self-Distillation, where an LLM conditioned on privileged info (correct answers or reasoning traces) supervises its weaker self via dense per-token feedback.

Original post →

More from Research

Research channel →