Apple Research: Uncovering the Effective Boundaries of On-Policy Distillation

Apple ML Research · rss · 2026-07-09

Apple ML Research published a post exploring when On-Policy Distillation is beneficial or harmful during the training of reasoning models.

The article notes that while this method provides dense supervisory signals per token, determining the optimal teacher model, the context for self-distillation, and whether these choices should vary by token typically requires expensive training. To address this, the research team proposed a training-free diagnostic method to reveal token-level dynamics.

Original post →

More from Research

Research channel →