Perplexity cuts tool-call failures 21.2% with hint-guided on-policy self-distillation
inductionheads · x · 2026-09-23
Perplexity shared new research on its multi-stage post-training pipeline: SFT, RL, RFT, and a final stage of on-policy self-distillation (OPSD). The motivation: RL environments are realistic but still lack the diversity and complexity of real production use cases, leaving a sim2real gap.
Key ideas of OPSD:
- Learn from problematic trajectories by constructing hints from tool-call errors and negative user feedback in follow-up conversations;
- Use self-distillation to fix those errors, combining RFT (forward KL/CE) for successful trajectories with OPSD (reverse KL) for problematic ones.
Applied to post-training their Computer model (demonstrated on GLM-5.2), a later checkpoint reduced tool-call failures by 21.2% relative to an earlier one in a live A/B test.
More from coding & agent
- Framer launches Skills: teach your design agent reusable workflows, design systems and CMS rules — soleio · 2026-09-23
- TesterArmy raises $1.2M pre-seed to build AI agents that test coding-agent-built apps — fernandorojo · 2026-09-23
- Shopify ditches React Native for native apps, and RedMonk says agents made rewrites feasible again — rseroter · 2026-09-23
- Two personal AIs tried to schedule coffee: agent interop needs a protocol — signulll · 2026-09-23
- Indie dev's SlideDev MCP gets first real user: screenshot to animated site via Claude — Separate-Topic4883 · 2026-09-23
- Dev open-sources awit: plain-Git agent work item tooling with graph-based work queue — nilreference · 2026-09-23