Microsoft's UniSteer Enables Robot to Learn Bead Tasks in 66 Mins
机器之心 · wechat · 2026-09-02
Microsoft Research proposes UniSteer, a method to enable efficient online fine-tuning of Vision-Language-Action (VLA) models on real robots by inverting human corrections into noise-space supervision. It addresses the incompatibility between human actions and the noise space in flow-matching architectures by freezing the pre-trained VLA and training only a lightweight noise actor. Experiments show UniSteer improves success rates from 20% to 90% across four tasks in an average of 66 minutes, completing a complex bead-placing task with just two full demonstration trajectories.
More from Embodied
- Digimon AI Project: Reward Hacking and Progress in PPO Training — redfoxkiller · 2026-09-02
- Qwen-Drive-1.0: A Vision-Language Foundation Model for Autonomous Driving — Qwen · 2026-09-02
- ZimaBlue: Evolving Generalizable World Action Models via Video Pre-training — JoyFutureAcademy · 2026-09-02
- Analyst report massively overestimates robot data generation — zephyr_z9 · 2026-09-02
- Markov Robotics demos sub-millimeter pick and place, highlighting zero-shot generalization for physical AGI — Scobleizer · 2026-09-02
- NVIDIA Warp hits 10M downloads; livestream to cover simulation and robotics workflows — milesmacklin · 2026-09-02