Microsoft's ShieldVLA Cuts VLA Safety Costs by 57% Using HJ Reachability
MicrosoftResearch · hf · 2026-09-22
Microsoft Research proposes ShieldVLA, a safety-aligned fine-tuning framework for Vision-Language-Action (VLA) models.
- Problem: existing methods rely on Lagrangian soft penalties on expected cumulative cost, causing residual violations or overly conservative behavior; dense per-step safety annotations are also scarce in visual domains.
- Method: builds on Hamilton-Jacobi (HJ) reachability, learning a model-free approximation of the reachability value function directly from visual observations to estimate the safe operating region. A learned safety critic gates policy optimization, separating reward maximization in feasible regions from recovery near unsafe states, avoiding persistent reward-cost trade-offs.
- Scalable supervision: rubric-based VLM safety scores convert semantic feedback into structured critic targets without manual cost labels.
- Results: across five navigation and manipulation benchmarks and multiple VLA backbones, ShieldVLA cuts cumulative safety cost by 57% on average and improves task success rate by +0.13 over SafeVLA.
More from Embodied
- Toronto's TTC Subway Cleaning Robots Get the 'Japan at Home' Meme Treatment — KadriJibraan · 2026-09-22
- Dev flips a Unitree G1 in a custom simulator, unsure if the result is realistic — rsasaki0109 · 2026-09-22
- VONDER launches camera-less AI glasses with bone-conduction mics, pre-orders from $299 — alex_verem · 2026-09-22
- Winning robotics: deploy early and lean on collaborative perception instead of chasing reliability — broodsugar · 2026-09-22
- Robot Astra installs new EM module on its own body, unprompted — mhmazur · 2026-09-22
- 50 promptable device concepts for letting agents inhabit everyday objects — genmon · 2026-09-22