Dream4ACT Unifies Video-Action Modeling Across Robot Embodiments, Hitting 89% on RoboTwin 2.0

Xiangyu Zhu · hf · 2026-10-05

Dream4ACT is a world model for joint video-action modeling across robot embodiments, solving the problem that joint-space action vectors lack image-space structure and vary across embodiments.

Results: 88.98% average success on RoboTwin 2.0 and 65.66 on TriWorldBench, supporting closed-loop manipulation.

Original post →

More from Embodied

Embodied channel →