PivotOPD: NVIDIA Paper Teaches Multi-Turn Agents to Recover From Early Pivotal Mistakes
MohitIyyer · x · 2026-10-01
New paper PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents (arXiv:2609.40285, authors include Jan Kautz and Ali Hatamizadeh of NVIDIA).
Motivation: with on-policy distillation (OPD) for language agents, one early incorrect action changes later states and errors compound across turns. Across three Qwen3 models (8B–235B), the authors find over half of failed rollouts contain a pivotal mistake — an action that moves the agent farther from the task — usually early; yet these remain recoverable by guiding the model for just a few turns afterward.
Method: PivotOPD jointly trains the student to prevent pivotal mistakes and recover from the states they create. At each pivotal mistake a teacher provides a gold action, then recovery actions for the next few turns: preventive distillation uses the gold action with reverse KL, while recovery distillation uses forward KL to transfer recovery behaviors the student rarely samples.
Results: against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students.
More from coding & agent
- Open-Source Coding Model IQuest-Q1 Hits Hugging Face, Works with Claude Code — ZabihullahAtal · 2026-10-01
- Opus 5.5 worked 23 hours straight to bring Omarchy's GPU desktop to WSL — sytelus · 2026-10-01
- Runway launches Visual Thinking: real-time video model that visualizes coding agent workflows — runwayml · 2026-10-01
- Devs now embrace alpha/beta software as coding agents make instability cheap — DavidKPiano · 2026-10-01
- Multi-harness RL guide: LFM2.5 jumps 42% to 54% with 31% fewer tool calls — SergioPaniego · 2026-10-01
- Claude Opus 5.5 Builds a Browser-Playable Rocket League Clone With Convincing Physics — minchoi · 2026-10-01