Tencent Open-Sources Two Embodied Foundation Models
机器之心 · wechat · 2026-07-15
Tencent RoboticsX, Fukuda Lab, and Tencent Hunyuan have jointly released and open-sourced two embodied foundation models: Hy-Embodied-VLM-1.0 and Hy-Embodied-RxBrain-1.0.
Hy-Embodied-VLM-1.0
- Focuses on the understanding-action-adaptation closed loop in embodied intelligence
- Covers physical state understanding, action-change reasoning, and temporal/adaptive reasoning
- Built on the latest Hunyuan A3B base, trained with large-scale embodied data
- Achieves performance comparable to the previous flagship A32B embodied model using only 1/10 of the compute across 37 evaluation tasks
Hy-Embodied-RxBrain-1.0
- Unified modeling of world understanding, reasoning/planning, and action consequence prediction
- Combines textual reasoning with visual imagination to provide dense multimodal conditions for downstream action models
- Training data includes 50,000+ hours of high-quality embodied data, comprising 210 million training samples
- Achieves a joint planning score of 0.68 on RxBrain-Bench, outperforming several comparative systems
Real-World Robots & Benchmarks
- Real-robot tasks cover multi-stage operations like arranging utensils, folding glasses, and throwing away trash
- Achieves an average success rate of 87% on the DOBOT X-Trainer and Fangzhou Wuxian A5
- The authors emphasize: visual imagination isn't just for "show," but serves as an intermediate cognitive representation actively utilized by action models
Related event: Tencent Hunyuan Open-Sources Two Embodied AI Foundation Models(3 posts)→
More from Embodied
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11