Dream4ACT Unifies Video-Action Modeling Across Robot Embodiments, Hitting 89% on RoboTwin 2.0
Xiangyu Zhu · hf · 2026-10-05
Dream4ACT is a world model for joint video-action modeling across robot embodiments, solving the problem that joint-space action vectors lack image-space structure and vary across embodiments.
- A shared visual action interface, "action views," renders target joint configurations from four virtual cameras via URDF forward kinematics, preserving embodiment-specific geometry while unifying representations.
- Observations and actions share one video autoencoder and diffusion transformer; masked flow-matching enables forward dynamics, inverse dynamics, and joint observation-action generation in a single model.
- A training-free, URDF-constrained multiview recovery mechanism extracts executable actions without embodiment-specific decoders.
Results: 88.98% average success on RoboTwin 2.0 and 65.66 on TriWorldBench, supporting closed-loop manipulation.
More from Embodied
- Munich Robotics Startup RobCo Hits $1B Valuation in Nine Months — lukas_m_ziegler · 2026-10-05
- Physical AI field session in Bengaluru to tackle post-deployment evals and retraining loops — carrycooldude · 2026-10-05
- Meta's NAVA-WAM pretrains robot action policies directly from action-free videos — meta · 2026-10-05
- Robotics' most-hyped model completed just 7 of 100 real manipulation tasks, demo reel hides the rest — carrycooldude · 2026-10-05
- Sergey Levine on robotics: hardware is good enough, the real gap is decision-making and data — 机器之心 · 2026-10-05
- SJTU's LIFT adds force sensing to VLAs with zero force-labeled pretraining data — jiqizhixin · 2026-10-05