Qwen details three embodied AI systems for navigation, manipulation and world models
青稞AI · wechat · 2026-07-29
Qwen’s robot team introduces three embodied-AI works: Qwen-RobotNav, Qwen-RobotManip, and Qwen-RobotWorld.
- Qwen-RobotNav unifies VLN, PointNav, ObjectNav, tracking, autonomous driving, and EQA into waypoint prediction, trained on 15.6M samples and reported as SOTA across several navigation benchmarks, including a zero-shot Unitree Go2 deployment.
- Qwen-RobotManip focuses on scalable cross-embodiment manipulation, trained on about 38,100 hours of manipulation data and said to rank first on multiple benchmarks, including RoboChallenge Table 30 v1.
- Qwen-RobotWorld is a language-conditioned embodied world model that predicts future visual states instead of low-level actions, trained on 8.6M video-text pairs, 200M+ frames, 20+ embodiments, and 500+ action categories.
The talk’s main message is “Alignment unlocks scaling”: embodied systems need explicit alignment across bodies, actions, cameras, history, instruction, and world transitions before scaling data and model size can translate into generalization and transfer.
More from Embodied
- LimX Demonstrates Dual-Arm Mobile Manipulation Robot for Pharmacy Tasks — chris_j_paxton · 2026-08-26
- Movie Scene Validates Facial ID Auth in Humanoid Robots — chrismatthieu · 2026-08-26
- No ChatGPT Moment for Robotics: Unitree Loses Nearly Half Its $66B Valuation — Sethwinterroth · 2026-08-26
- Developer tries running neural models on a mosquito detector — yacineMTB · 2026-08-26
- Robot dancing is getting insanely good, smooth moves approaching human level — drgoldenpants · 2026-08-26
- Technology and AI accelerate agriculture and support farmers — DynamicWebPaige · 2026-08-26