Qwen Launches Unified Embodied Control Model

CyberRobooo · x · 2026-07-19

Alibaba's Tongyi Qwen has introduced **Qwen-VLA**, a unified Vision-Language-Action model designed to directly control robots of various forms, including humanoid and dual-arm setups. It integrates manipulation, navigation, trajectory prediction, and cross-form control into a single system, using "embodied perception prompts" to adapt to different robot embodiments without needing to train a separate control head for each form. The post also claims that Qwen-VLA matches or surpasses specialized models on key benchmarks, achieving a **76.9%** Out-of-Distribution (OOD) success rate on real-world ALOHA bimanual tasks.

Original post →

More from Embodied

Embodied channel →