Xiaomi Open-Sources VLA Robot Foundation Model
XiaomiRobotics · hf · 2026-07-20
Xiaomi Robotics released Xiaomi-Robotics-1, a Vision-Language-Action (VLA) foundation model designed to perform various mobile manipulation tasks directly in unseen environments and adapt to new tasks with minimal fine-tuning.
The paper employs a two-stage training approach:
- Pre-training: Utilizes over 100,000 hours of real-world manipulation trajectories collected via UMI devices, which are converted into natural language state-transition descriptions using an automated annotation pipeline.
- Post-training: Aligns the model more closely with specific robot embodiments and common human instruction formats.
Experiments show the model scales effectively with data volume and model size, outperforming existing methods on multiple simulation benchmarks. It achieves a 57.6% success rate on RoboCasa365 (beating the previous 46.6%) and scores 20.07 on RoboDojo (up from 13.07). Code and model weights will be released.
Related event: Xiaomi Open-Sources VLA Robotics Foundation Model(3 posts)→
More from Embodied
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11