NVIDIA paper: verify-then-execute lifts TerminalBench Pass@1 from 50.0% to 68.0%
dair_ai · x · 2026-10-04
An NVIDIA paper on test-time compute for terminal agents finds you should sample several candidate shell commands and verify them before running one — and spend more on the verifier than on extra samples.
Key results:
- With a GPT-5.6 Sol verifier choosing among 8 sampled actions, TerminalBench-Lite Pass@1 rises from 50.0% to 68.0%
- With a weak verifier, extra samples add almost nothing
Mid-Harness leaves the generator and harness unchanged and works between them. When the small TMAX-9B model verifies its own candidates, pairwise comparison works best, and distilling the strong verifier in helps further. Combining action sampling with trajectory sampling reaches higher success at lower token cost than sampling full trajectories alone.
More from coding & agent
- Hosted chat UI for OpenAI's Agents API pitched: plug in agents, no FE code — DarasStayHome · 2026-10-04
- Alloyqa lets coding agents work in the cloud and hand back a PR to review — AdventurousAge8767 · 2026-10-04
- OpenAI releases 34-page whitepaper on how it builds AI agents — mdancho84 · 2026-10-04
- Meta engineer: an AI session fixes top JS errors daily and pings owners — Vjeux · 2026-10-04
- 19-year-old builds Telegram AI agent: 'It's just a wrapper, that's all' — PowerOk7047 · 2026-10-04
- Monitoring Claude/Codex all day is frying my dopamine system, dev says — ethanCaballero · 2026-10-04