Xiaomi Robotics Trains VLA Model With 100k Hours of Human Handheld Data
TinfoilTricorn · x · 2026-08-05
Xiaomi Robotics recently showcased new progress with its Vision-Language-Action (VLA) model. The model was pretrained on over 100,000 hours of real-world UMI handheld-gripper trajectories.
This approach attempts to replace costly robot demonstrations with low-cost, human-recorded manipulation data. After robot alignment, it can handle tasks like shoe storage, luggage packing, and table organization in unseen rooms. In controlled tests, fine-tuning for less than 10 hours per task yielded an average 75% success rate across phone packing, printer refilling, and laundry loading.
Related event: Xiaomi Open-Sources Robotics Foundation Model Robotics-1(4 posts)→
More from Embodied
- DOBOT Unveils LUMO: A 1.3-Meter Home Humanoid Focused on Social Intelligence — ChrisGPT · 2026-08-06
- FCC Advanced Robot Ban: Devices with >35% Foreign Components Face Prohibition — mattfreed · 2026-08-06
- Physical Robot Chess Champions Outperform Humans Like Stockfish Beating Bots — srchvrs · 2026-08-06
- Embodied AI in Action: Robot Autonomously Collects Multi-Lighting Image Data Overnight — alexisgallagher · 2026-08-05
- Ginkgo to Build 4 New Autonomous Labs at MIT, Caltech, and More — ProfBuehlerMIT · 2026-08-05
- Jensen Huang on Digital Immortality: Sending His Humanoid Robot to Space — r0ck3t23 · 2026-08-05