MSRA, UTS and Tsinghua unveil UniSteer: injecting human corrections into RL for flow-matching VLAs
jiqizhixin · x · 2026-09-11
Microsoft Research Asia, University of Technology Sydney, and Tsinghua University present UniSteer, tackling the interface gap between human guidance and lightweight RL for flow-matching VLA policies.
- Real-robot RL is too expensive for blind exploration, so human guidance is the natural fix — but flow-matching VLAs optimize initial noise, and human actions have no direct mapping into that noise space.
- UniSteer bridges this via approximate action-to-noise inversion: when a human takes over and corrects the robot, the correction is inverted into target noise.
- The data flows into two buffers simultaneously: a demonstration buffer for direct MSE supervision of the actor, and an RL buffer that helps the critic learn these high-value state-noise pairs. Both supervised and reinforcement learning update the same lightweight noise actor while the pretrained VLA backbone stays frozen.
More from Embodied
- New joystick control mode demoed for Masked Mimic robot policy — carlosdponx · 2026-09-11
- OpenAI's Astra shows in-context learning for mobile robot manipulation from video alone — npew · 2026-09-11
- Actor Labs and Physical Intelligence to host 36-hour physical AI hackathon in Mountain View — Darpinian · 2026-09-11
- Household robot remote assistance raises a privacy question: who's watching the cameras? — VraserX · 2026-09-11
- NVIDIA Details Full-Stack Robotaxi Platform as Market Heads Toward $400B by 2035 — nvidia · 2026-09-11
- Tesla Cybercab carefully passes cyclist on narrow hillside road in viral clip — EricETesla · 2026-09-11