Being-M0.7: Humanoid Robot Whole-Body Action Model
机器之心 · wechat · 2026-07-15
A detailed Synced review introduces Being-M0.7 by Zhizai WuJie: a LatentWorld-Action Model designed for whole-body mobile manipulation in humanoid robots.
Key Conclusions
- This is the world's first LatentWorld-Action Model for humanoid robot whole-body mobile manipulation.
- The goal isn't just tabletop dexterity, but enabling robots to simultaneously learn "where to go, how to turn, and how to coordinate limbs."
- The model successfully performed high-difficulty tasks on real humanoid robots, such as liquid interaction, mirror retrieval, long-horizon transport, and obstacle avoidance.
Methods and Data
- Employs a Vision-Motion MoT architecture to separately process vision and motion modalities, interacting via shared attention.
- Pre-trained on over 10,000 hours of data, including human first-person videos, video-motion paired data, and pure human motion sequences.
- Constructed a unified motion representation shared by humans and humanoid robots to bridge morphological differences.
- Learns future state changes and motion trajectories through flow matching objectives.
- Real-world data collection used a PICO VR-based whole-body teleoperation system, with a motion expert translating high-level predictions into executable control commands.
Four Real-World Demos
- Catching Fish: Tool-based catching under fluid dynamics and visual distortion.
- Mirror Retrieval: Inferring the position of hidden objects using mirror reflections.
- Mobile Object Placement: Continuously executing walking, grasping, transferring, and turning.
- Box Carrying & Obstacle Avoidance: Passing sideways through obstacles under load and in narrow spaces.
The article emphasizes that the competitive focus of embodied AI is shifting from the hardware body to model capabilities and scalable data flywheels.
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11