NVIDIA's PivotOPD teaches agents to prevent and recover from pivotal early mistakes
NVIDIAAI · x · 2026-10-08
NVIDIA researchers present PivotOPD, an on-policy distillation framework for multi-turn language agents. Their finding: across three Qwen3 models (8B–235B), over half of failed rollouts hinge on a single early 'pivotal mistake', usually recoverable within a few turns. During training, a teacher model supplies a gold action at each pivotal mistake plus a recovery action for subsequent steps, jointly teaching the student to prevent such errors and to recover from the states they create. PivotOPD achieves the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA, with gains transferring to SWE-Bench Verified. Code coming soon.
More from coding & agent
- Agents love tidying up files nobody asked them to touch — JFPuget · 2026-10-08
- Open-source bridge turns a browser chat tab into an OpenAI-compatible API endpoint — harshanacz · 2026-10-08
- GraphRAG Bug: Deleted Documents Stay Indexed and Retrievable After Updates — JeremyCMorgan · 2026-10-08
- Paper: Vibe Coding Kills Open Source as AI-Recommended Repos Lose Stars — soumitrashukla9 · 2026-10-08
- Every publishes definitive guide to Compound Engineering, the philosophy behind 7k-star plugin — every · 2026-10-08
- Claude Haiku 5.5 matches GPT-6 Luna pricing but a stingier tokenizer hides a 1.25x cost hike — Simon Willison · 2026-10-08