Xiaomi Open-Sources VLA Robot Foundation Model

XiaomiRobotics · hf · 2026-07-20

Xiaomi Robotics released Xiaomi-Robotics-1, a Vision-Language-Action (VLA) foundation model designed to perform various mobile manipulation tasks directly in unseen environments and adapt to new tasks with minimal fine-tuning.

The paper employs a two-stage training approach:

Experiments show the model scales effectively with data volume and model size, outperforming existing methods on multiple simulation benchmarks. It achieves a 57.6% success rate on RoboCasa365 (beating the previous 46.6%) and scores 20.07 on RoboDojo (up from 13.07). Code and model weights will be released.

Related event: Xiaomi Open-Sources VLA Robotics Foundation Model(3 posts)→

Original post →

More from Embodied

Embodied channel →