NVIDIA paper: verify-then-execute lifts TerminalBench Pass@1 from 50.0% to 68.0%

dair_ai · x · 2026-10-04

An NVIDIA paper on test-time compute for terminal agents finds you should sample several candidate shell commands and verify them before running one — and spend more on the verifier than on extra samples.

Key results:

Mid-Harness leaves the generator and harness unchanged and works between them. When the small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier in helps further. Combining action sampling with trajectory sampling reaches higher success at lower token cost than sampling full trajectories alone.

Original post →

More from coding & agent

coding & agent channel →