NVIDIA's PivotOPD Teaches Agents to Prevent and Recover From Pivotal Mistakes

nvidia · hf · 2026-10-01

PivotOPD is an on-policy distillation framework that trains agents both to avoid pivotal mistakes (found in over half of failed rollouts across Qwen3 8B-235B, usually early) and to recover from the states they create. It beats 13 baselines, gains +5.5% on ALFWorld with the 1.7B student, and lifts a Nemotron-3.5 student's SWE-Bench Verified resolve rate by 3.2%.

Original post →

More from coding & agent

coding & agent channel →