NVIDIA's VERA scales 9,000+ verifiable environments for agent self-improvement
cihangxie · x · 2026-10-07
NVIDIA Research and university collaborators released VERA, a system of verifiable environments for agentic RSI where agents improve both model weights and harness skills.
- 9,000+ verifiable long-horizon environments with restartable sandboxes, evidence-based rubrics, and executable checks
- Model–harness co-evolution: alternating model training and skill updates guided by failure reports and verifiers
- VERA-CoWork-27B hits 43.8% on Automation-Bench (+19.2 pts over Qwen3.8-27B), ranking second ahead of Claude Opus; VERA-Med-27B reaches 76.4% on AgentClinic (+32.5 pts)
- Estimated $240k for the MedResearch RSI pipeline, 54.7% of it spent on environment preparation — the bottleneck is environments, not learning algorithms
- Models, data, paper, and blog are open-sourced
Related event: NVIDIA Unveils VERA for Co-Evolving Agents and Harnesses(2 posts)→
More from coding & agent
- Claude Haiku 5.5 lands in Cursor at 10x cheaper for short requests; Sonnet 5.5 cache reads cut 50% — mattyp · 2026-10-08
- LangChain ships Managed Deep Agents v0.9 with agent-created schedules and per-run config — LangChain · 2026-10-08
- Vercel's AI SDK hits 30 million weekly downloads, up 6x in under a year — lgrammel · 2026-10-08
- GitHub HydraFusion adds local model routing as Microsoft ships MAI Code 1.1 at 3-bit, 256K — BenBajarin · 2026-10-08
- Devin Mobile hands-on: cloud agents running on Linux, macOS, or Windows from anywhere — DevinAI · 2026-10-08
- GodTerm: open-source tool pools multiple Claude Code and Grok accounts with limit rollover — Daniel_Farinax · 2026-10-08