PivotOPD: NVIDIA Paper Teaches Multi-Turn Agents to Recover From Early Pivotal Mistakes

MohitIyyer · x · 2026-10-01

New paper PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents (arXiv:2609.40285, authors include Jan Kautz and Ali Hatamizadeh of NVIDIA).

Motivation: with on-policy distillation (OPD) for language agents, one early incorrect action changes later states and errors compound across turns. Across three Qwen3 models (8B–235B), the authors find over half of failed rollouts contain a pivotal mistake — an action that moves the agent farther from the task — usually early; yet these remain recoverable by guiding the model for just a few turns afterward.

Method: PivotOPD jointly trains the student to prevent pivotal mistakes and recover from the states they create. At each pivotal mistake a teacher provides a gold action, then recovery actions for the next few turns: preventive distillation uses the gold action with reverse KL, while recovery distillation uses forward KL to transfer recovery behaviors the student rarely samples.

Results: against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students.

Related event: NVIDIA's PivotOPD teaches multi-turn agents to recover from pivotal mistakes(2 posts)→

Original post →

More from coding & agent

coding & agent channel →