Xiaomi Robotics Trains VLA Model With 100k Hours of Human Handheld Data

TinfoilTricorn · x · 2026-08-05

Xiaomi Robotics recently showcased new progress with its Vision-Language-Action (VLA) model. The model was pretrained on over 100,000 hours of real-world UMI handheld-gripper trajectories.

This approach attempts to replace costly robot demonstrations with low-cost, human-recorded manipulation data. After robot alignment, it can handle tasks like shoe storage, luggage packing, and table organization in unseen rooms. In controlled tests, fine-tuning for less than 10 hours per task yielded an average 75% success rate across phone packing, printer refilling, and laundry loading.

Related event: Xiaomi Open-Sources Robotics Foundation Model Robotics-1(4 posts)→

Original post →

More from Embodied

Embodied channel →