HiDream launches embodied world model HiDream-O1-Embodied, tops RoboColiseum robustness board at 0.692

机器之心 · wechat · 2026-09-07

HiDream.ai released HiDream-O1-Embodied, an embodied world model aimed at turning video generation into action control with stronger physical perception and dynamic prediction. On its first RoboColiseum evaluation, the model topped the Robustness sub-board with an average score of 0.692. The platform runs 78 high-fidelity simulation tasks, testing stability under changed lighting, materials, camera positions and rewritten instructions.

Key technical claims include intent-level language understanding beyond keyword matching, multi-view visual perception that keeps running when some views fail, and fault-tolerant training with deliberately degraded conditions. Data-wise, HiDream uses a "real base + generative augmentation" paradigm, expanding mocap data 100x while preserving physical constraints.

CTO Yao Ting said full-modality representation, causal reasoning and physical world modeling together form a complete world model foundation. A month earlier HiDream topped the WBench Navi board at 80.9 with its interactive world model HiDream-O1-World.

Original post →

More from Embodied

Embodied channel →