Microsoft's DiVeR Reweights VLA Verifier Learning Toward Decision-Critical States

MicrosoftResearch · hf · 2026-10-07

Microsoft Research introduces DiVeR, improving verifier-guided test-time scaling for Vision-Language-Action (VLA) policies. Existing classification verifiers treat all visited states equally, though only sparse decision-critical states offer meaningful action discrimination.

Method

Results: Consistent task success improvements on LIBERO, RoboCasa, and real-world Franka Research 3 experiments, with negligible verifier inference overhead.

Original post →

More from Embodied

Embodied channel →