NVIDIA paper: a better judge lifts terminal agent success from 50% to 68% without retraining

rohanpaul_ai · x · 2026-10-02

NVIDIA researchers argue you should give your terminal agent a better judge, not more options. A small model drafts 8 candidate commands per step and a stronger frontier-model verifier picks one before execution, raising success rate from 50% to 68% with no retraining. When the small model judged its own drafts, gains were much smaller — more options only help if the judge can tell them apart.

Related event: NVIDIA Paper: Adding a Verifier Lifts Terminal Agent Success to 68%(3 posts)→

Original post →

More from coding & agent

coding & agent channel →