Perplexity details hint-guided self-distillation: agent tool-call failures cut from 2.24% to 1.77%
perplexity_ai · x · 2026-09-23
Perplexity shared its agent training recipe: rejection sampling fine-tuning imitates useful steps from successful sessions, while hint-guided self-distillation corrects errors—having GLM 5.2 score the same turn with and without a corrective hint and aligning the hint-free predictions. The annotation pipeline traces feedback to the responsible decision and checks hints against pre-mistake information to reduce hindsight bias. Results: with hints, the unchanged model avoided the original failure in 93.7% of cases (up from 75.1%), and live tool-call failures fell from 2.24% to 1.77% between trained versions without inference-time hints. Training starts with RL in synthetic environments, then real-world sessions; PII and opt-out sessions are excluded.
More from coding & agent
- TesterArmy raises $1.2M pre-seed to build AI agents that test coding-agent-built apps — fernandorojo · 2026-09-23
- Anthropic launches Claude Opus 5.5: matches Fable 5.1 on most tasks at 40% lower cost — jyangballin · 2026-09-23
- Shopify ditches React Native for native apps, and RedMonk says agents made rewrites feasible again — rseroter · 2026-09-23
- Hamel Husain & Shreya Shankar Share Their Playbook for Building AI Eval Systems — HamelHusain · 2026-09-23
- Two personal AIs tried to schedule coffee: agent interop needs a protocol — signulll · 2026-09-23
- Indie dev's SlideDev MCP gets first real user: screenshot to animated site via Claude — Separate-Topic4883 · 2026-09-23