NVIDIA's Mid-Harness: a strong verifier boosts terminal agent Pass@1 from 50% to 68% on TerminalBench-Lite
rohanpaul_ai · x · 2026-10-02
NVIDIA's new paper introduces Mid-Harness, which samples and verifies candidate actions at the model-harness boundary before execution, leaving generator and harness unchanged. With a TMAX-9B generator, a GPT-5.6 Sol verifier raises Pass@1 on TerminalBench-Lite from 50.00% to 68.03% using 8 sampled actions; weak verifiers gain little from more sampling. Pairwise verification works best for self-verification, and distilling the stronger verifier's responses back into TMAX-9B further improves Pass@1.
Related event: NVIDIA Paper: Adding a Verifier Lifts Terminal Agent Success to 68%(3 posts)→
More from coding & agent
- Microsoft paper: coding agent optimizing prompts from logs beats GEPA at ~$1.60 — rohanpaul_ai · 2026-10-02
- Stop reaching for the biggest model: a cost-efficient Cursor/Codex setup with GPT-6.1 Sol — chiliraupe · 2026-10-02
- Dot isn't better than Codex or Claude Code — it's a different, undervalued take on OpenClaw/Hermes — gabrielchua · 2026-10-02
- Engineering.com parent Arrowfly launches year-round AI for Engineers initiative amid vibe-CAD era — burhop · 2026-10-02
- Run Fewer Agents: exe.dev argues task management is a band-aid, proposes fast models for human comms — sull · 2026-10-02
- Developer gives an AI agent $1,000 to run a live-streamed hedge fund — kleffew94 · 2026-10-02