NVIDIA's PivotOPD teaches agents to prevent and recover from pivotal early mistakes

NVIDIAAI · x · 2026-10-08

NVIDIA researchers present PivotOPD, an on-policy distillation framework for multi-turn language agents. Their finding: across three Qwen3 models (8B–235B), over half of failed rollouts hinge on a single early 'pivotal mistake', usually recoverable within a few turns. During training, a teacher model supplies a gold action at each pivotal mistake plus a recovery action for subsequent steps, jointly teaching the student to prevent such errors and to recover from the states they create. PivotOPD achieves the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA, with gains transferring to SWE-Bench Verified. Code coming soon.

Original post →

More from coding & agent

coding & agent channel →