Microsoft's DiVeR Reweights VLA Verifier Learning Toward Decision-Critical States
MicrosoftResearch · hf · 2026-10-07
Microsoft Research introduces DiVeR, improving verifier-guided test-time scaling for Vision-Language-Action (VLA) policies. Existing classification verifiers treat all visited states equally, though only sparse decision-critical states offer meaningful action discrimination.
Method
- Estimates decision criticality from the dispersion of sampled action representations
- Reweights verifier learning toward states where action selection matters most, requiring no step-level annotations or extra environment interaction
Results: Consistent task success improvements on LIBERO, RoboCasa, and real-world Franka Research 3 experiments, with negligible verifier inference overhead.
More from Embodied
- Boston Dynamics Names Former Amazon Executive Rohit Prasad as CEO — SumitGup · 2026-10-07
- Designing a robotics lab in 20 minutes: parametric 3D layout from room specs and robot dimensions — Stefania_druga · 2026-10-07
- Dev builds a spatial AR interface for the Dobot Rover X1 robot in Godot on Quest 3 — Scobleizer · 2026-10-07
- 7 minutes per motor, ~$5 labor cost: why robotics automation is the only path for Western manufacturing — avlok · 2026-10-07
- Reward-DAgger: generalist reward models enable task-agnostic runtime monitoring for robots — ebiyik_ · 2026-10-07
- NVIDIA's VeriFine Co-Evolves Policy and Judge to Scale Self-Improvement in Embodied Reasoning — nvidia · 2026-10-07