Perplexity cuts tool-call failures 21.2% with hint-guided on-policy self-distillation

inductionheads · x · 2026-09-23

Perplexity shared new research on its multi-stage post-training pipeline: SFT, RL, RFT, and a final stage of on-policy self-distillation (OPSD). The motivation: RL environments are realistic but still lack the diversity and complexity of real production use cases, leaving a sim2real gap.

Key ideas of OPSD:

Applied to post-training their Computer model (demonstrated on GLM-5.2), a later checkpoint reduced tool-call failures by 21.2% relative to an earlier one in a live A/B test.

Related event: Perplexity Details Hint-Guided Self-Distillation for Training Its Computer Agent(5 posts)→

Original post →

More from coding & agent

coding & agent channel →