NVIDIA's VERA co-evolves agent harness with model weights, 9B agent beats baselines by 10+ points
omarsar0 · x · 2026-10-07
A new NVIDIA paper, VERA, demonstrates co-evolving the agent harness alongside model weights — a trend the author says frontier labs have already adopted to own the intelligence stack:
- Converts benchmark trajectories into 9,000+ restartable sandboxes with rubric scoring, keeping only environments that run and can be scored from observable evidence.
- Updates both weights and harness: a harness edit is kept only if it passes self-tests and gains at least 5 points on the dev set; model checkpoints dropping more than 20% are rejected.
- Results: the co-evolved 9B agent beats the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 on AutoMedBench.
Takeaway: building verifiable environments for agents and training the harness in the loop is becoming standard practice at frontier companies.
Related event: NVIDIA Unveils VERA for Co-Evolving Agents and Harnesses(2 posts)→
More from coding & agent
- LangChain engineer built an ACP coding agent that replaced Claude Code for 9 months — Hacubu · 2026-10-08
- 30 Real Business Workflow Tests: Keep Agent Evaluation Simple — VibeMarketer_ · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Long-running agents suffer 'constraint amplification': a subtle form of context rot — generativist · 2026-10-08
- Nautilo Ships Text+Vision Model Split, Preps Price/Security-Based Model Routing Gateway — Dan_Jeffries1 · 2026-10-08
- A fine-tuned 9B beats a 31B model: 600 labels, $0.12, 91% accuracy — julsimon · 2026-10-08