NVIDIA's Mid-Harness Scales Actions at the Model-Harness Boundary, Lifting TerminalBench Pass@1 to 68.03%
nvidia · hf · 2026-10-01
NVIDIA's Mid-Harness samples and verifies candidate actions before execution to improve terminal-agent reliability. With TMAX-9B, a GPT-5.6 Sol verifier raises TerminalBench-Lite Pass@1 from 50.00% to 68.03% with 8 sampled actions; pairwise verification works best when TMAX-9B verifies itself, and distilling verifier responses back in helps further. Combining action and trajectory scaling beats generating more trajectories at lower token cost, generalizing across models, benchmarks, and harnesses.
More from coding & agent
- Creator Shares His Full Daily Claude Skills Library: /grill-me, /humanizer and More — rubenhassid · 2026-10-01
- Google Cloud Launches Advent of Agents Season 3: 31 Days of Free Production Agent Tutorials — Saboo_Shubham_ · 2026-10-01
- Multiple AI agents per person is going normal: OpenClaw now runs on a $10/mo VPS — steipete · 2026-10-01
- Early user: GPT-6 Sol feels slower and dumber than 5.6 in Codex — burkov · 2026-10-01
- No Evals, Building Blind: Contextual Embedding Models Must Rank Disambiguating Chunks — antoine_chaffin · 2026-10-01
- Jev Sentinel open-sources per-action agent monitor that scored 53,870 HF payloads, flagging 98.3% — schwentker · 2026-10-01