Alibaba DAMO Academy's 4D World Model

机器之心 · wechat · 2026-07-17

Alibaba DAMO Academy proposed **RynnWorld-4D** to address the representational limitations of traditional 2D video world models in robotic manipulation: pure pixel prediction lacks depth, motion fields, and physical constraints, making precise 3D operations difficult. This method simultaneously generates **RGB, Depth, and Flow** within a single diffusion framework, forming a 4D prediction using explicit geometric and motion information. Methodologically, it employs a three-branch Transformer to model texture, geometry, and motion respectively, aligning them via cross-modal attention and frame-level 3D positional encoding. Training is split into three phases, incorporating a large-scale data pipeline with 2.54 亿帧 of 4D annotated data. For policy execution, the model extracts 4D information directly from intermediate features, outputting actions in a single forward pass—completing inference in about 1.1 seconds and running at roughly 9Hz on real robots. Experiments show that this 4D representation outperforms baselines in depth, optical flow, and robotic bimanual dexterous manipulation tasks, yielding significant gains in spatially precise tasks like handovers, lid placement, and bowl stacking. The article concludes that the next bottlenecks lie in inference acceleration and multi-view expansion.

Related event: Alibaba DAMO Academy Releases RynnWorld-4D for Robotics(2 posts)→

Original post →

More from Embodied

Embodied channel →