NVIDIA's PivotOPD Teaches Agents to Prevent and Recover From Pivotal Mistakes
nvidia · hf · 2026-10-01
PivotOPD is an on-policy distillation framework that trains agents both to avoid pivotal mistakes (found in over half of failed rollouts across Qwen3 8B-235B, usually early) and to recover from the states they create. It beats 13 baselines, gains +5.5% on ALFWorld with the 1.7B student, and lifts a Nemotron-3.5 student's SWE-Bench Verified resolve rate by 3.2%.
More from coding & agent
- Priors: An Onchain Credit Bureau for AI Agents Emerges — econoar · 2026-10-01
- Tripo API integrates with Unbound for real-time sculpting of AI-generated 3D assets — yshan2u · 2026-10-01
- Veteran game dev lists 10 ways AI saves time: crash logs, CMake, 400ms frame hitches — draginol · 2026-10-01
- NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03% — nvidia · 2026-10-01
- Amazon's SMART Self-Evolving Multi-Agent System Tops All 15 Subtitle Arena Directions, Cuts Penalty 6.9% — amazon · 2026-10-01
- Gary Bernhardt hits all-time low faith in AI agents: they "fix" tests by deleting them — sidjustice_ · 2026-10-01