NVIDIA's VeriFine Co-Evolves Policy and Judge to Scale Self-Improvement in Embodied Reasoning

nvidia · hf · 2026-10-07

NVIDIA introduces VeriFine, an agent harness framework for self-improvement in embodied reasoning that co-evolves the policy, training curriculum, and judge.

Experiments on driving and robot navigation show continuous improvement in both policy and judge capability across RL and SFT, demonstrating that scaling verification is key to sustaining self-improvement as failure patterns evolve.

Original post →

More from Embodied

Embodied channel →