Perplexity's post-training method cuts agent tool call failures by 21%
beffjezos · x · 2026-09-23
- Perplexity shared new research on post-training its Computer agent: learning from real user sessions by imitating good trajectories and explicitly correcting avoidable mistakes like bad tool calls, even when the overall trajectory succeeded.
- The method combines rejection sampling fine-tuning (RFT) with hint-guided self-distillation.
- Live A/B tests show tool call failures cut by 21%.
- beffjezos amplified it, arguing there's "huge alpha" in post-training custom models that too few people are talking about.
More from coding & agent
- Stripe adds WebMCP to Checkout: 42% fewer tokens, 38% fewer tool calls for agents — gaganghotra_ · 2026-09-23
- Developer says delegating all coding to his AI agent has him learning and experimenting faster than ever — tristanbob · 2026-09-23
- Matt Shumer credits community demo, restarts NYC Open World loop with Opus 5.5 — mattshumer_ · 2026-09-23
- Jeffrey Emanuel builds a Claude Code Skill to port Rust projects to Bend 2 — doodlestein · 2026-09-23
- Dev builds playable Game Boy Color with Claude Opus 5.5 that runs Super Mario — chrisfirst · 2026-09-23
- Open-source Unreal Agent claims 39% cheaper than Codex+Astra on Terminal-Bench 4.0 — Hesamation · 2026-09-23