NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03%

nvidia · hf · 2026-10-01

NVIDIA's Mid-Harness samples and verifies candidate actions before execution to improve terminal-agent reliability. With TMAX-9B, a GPT-5.6 Sol verifier raises TerminalBench-Lite Pass@1 from 50.00% to 68.03% with 8 sampled actions; pairwise verification works best when TMAX-9B verifies itself, and distilling verifier responses back in helps further. Combining action and trajectory scaling beats generating more trajectories at lower token cost, generalizing across models, benchmarks, and harnesses.

Original post →

More from coding & agent

coding & agent channel →