Perplexity details hint-guided self-distillation: agent tool-call failures cut from 2.24% to 1.77%

perplexity_ai · x · 2026-09-23

Perplexity shared its agent training recipe: rejection sampling fine-tuning imitates useful steps from successful sessions, while hint-guided self-distillation corrects errors—having GLM 5.2 score the same turn with and without a corrective hint and aligning the hint-free predictions. The annotation pipeline traces feedback to the responsible decision and checks hints against pre-mistake information to reduce hindsight bias. Results: with hints, the unchanged model avoided the original failure in 93.7% of cases (up from 75.1%), and live tool-call failures fell from 2.24% to 1.77% between trained versions without inference-time hints. Training starts with RL in synthetic environments, then real-world sessions; PII and opt-out sessions are excluded.

Related event: Perplexity Details Hint-Guided Self-Distillation for Training Its Computer Agent(5 posts)→

Original post →

More from coding & agent

coding & agent channel →