NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite

rohanpaul_ai · x · 2026-10-02

NVIDIA's new paper introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before execution, leaving generator and harness unchanged. With a TMAX-9B generator, a GPT-5.6 Sol verifier raises Pass@1 on TerminalBench-Lite from 50.00% to 68.03% using 8 sampled actions; weak verifiers gain little from more sampling. Pairwise verification works best for self-verification, and distilling the stronger verifier's responses back into TMAX-9B further improves Pass@1.

Related event: NVIDIA Paper: Adding a Verifier Lifts Terminal Agent Success to 68%(3 posts)→

Original post →

More from coding & agent

coding & agent channel →