Robot world model trained on 15 hours of video generalizes to unseen bodies
jon_barron · x · 2026-07-24
Researchers trained a video world model using only 15 hours of video from a single-arm robot.
- It generalizes zero-shot to unseen embodiments, including orangutans.
- The model also learned an emergent capability not explicitly trained for: given a desired object motion, it can synthesize the robot motion that would produce it.
- The key idea is to represent actions as masked video.
The post highlights a compact but surprisingly general robot-world-model setup, with an emphasis on embodiment transfer and action generation.
More from Embodied
- Augmented reality CAD turns a hardware assembly into a maker joke — _Stocko_ · 2026-07-24
- VTAP Gripper aims to bridge the gap between simple grippers and dexterous robot hands — YunzhuLiYZ · 2026-07-24
- DoorDash says its robotics and drone bets started eight years ago — garrytan · 2026-07-24
- HP and AMD pitch AI workstations for local models and multi-agent workflows — gaganghotra_ · 2026-07-24
- Embodied AI needs years of expert trade knowledge to learn real-world constraints — Exp_Mark · 2026-07-24
- Tesla robotaxi rider says Tampa trip felt smooth and near inflection point — downingARK · 2026-07-24