Xiaomi Open-Sources VLA Robot Foundation Model
XiaomiRobotics · hf · 2026-07-20
Xiaomi Robotics released Xiaomi-Robotics-1, a Vision-Language-Action (VLA) foundation model designed to perform various mobile manipulation tasks directly in unseen environments and adapt to new tasks with minimal fine-tuning.
The paper employs a two-stage training approach:
- Pre-training: Utilizes over 100,000 hours of real-world manipulation trajectories collected via UMI devices, which are converted into natural language state-transition descriptions using an automated annotation pipeline.
- Post-training: Aligns the model more closely with specific robot embodiments and common human instruction formats.
Experiments show the model scales effectively with data volume and model size, outperforming existing methods on multiple simulation benchmarks. It achieves a 57.6% success rate on RoboCasa365 (beating the previous 46.6%) and scores 20.07 on RoboDojo (up from 13.07). Code and model weights will be released.
Related event: Xiaomi Open-Sources VLA Robotics Foundation Model(3 posts)→
More from Embodied
- Sunday Robotics Hires Strategic Projects Lead to Accelerate Home Robots — tonyzzhao · 2026-07-22
- Tesla’s summer update lets Grok make calls, control climate, and play music — Polymarket · 2026-07-22
- Hugging Face Teases Agentic Training Environments with OpenEnv for August Launch — mervenoyann · 2026-07-22
- NVIDIA shows 22 SIGGRAPH papers and Omniverse tools for robot simulation — facontidavide · 2026-07-22
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22