NVIDIA's VeriFine Co-Evolves Policy and Judge to Scale Self-Improvement in Embodied Reasoning
nvidia · hf · 2026-10-07
NVIDIA introduces VeriFine, an agent harness framework for self-improvement in embodied reasoning that co-evolves the policy, training curriculum, and judge.
- Policy Improvement Loop: a rubric judge diagnoses recurring failure patterns, builds an adaptive curriculum, and optimizes the policy.
- Judge Improvement Loop: when progress plateaus and verification becomes the bottleneck, humans are selectively queried on informative failures, and coactive calibration aligns human-agent disagreements toward an objective rubric of physical reasoning.
Experiments on driving and robot navigation show continuous improvement in both policy and judge capability across RL and SFT, demonstrating that scaling verification is key to sustaining self-improvement as failure patterns evolve.
More from Embodied
- Minerva Humanoids debuts Roger robot for dangerous jobs, raises $10M pre-seed — Scobleizer · 2026-10-07
- 7 minutes per motor, ~$5 labor cost: why robotics automation is the only path for Western manufacturing — avlok · 2026-10-07
- Reward-DAgger: generalist reward models enable task-agnostic runtime monitoring for robots — ebiyik_ · 2026-10-07
- EmbodiedSmith: Recursive Self-Improvement Flywheel Scales Embodied Training Data in Simulation — Yikai Qin · 2026-10-07
- ETH student documents building and calibrating a UMI gripper from scratch — SongShuran · 2026-10-07
- Kangwook Lee on how games can help build AI for the physical world — Kangwook_Lee · 2026-10-07