Zero-Shot Robot World Model Trained with 15 Hours of Video
Using only 15 hours of single-arm robot video, researchers trained a video world model via Masked Visual Actions (MVA). The model demonstrates impressive zero-shot generalization, adapting to unseen embodiments and even orangutans.
2026-07-24 ~ 2026-07-25 · 3 related posts
- Robot world model trained on 15 hours of video generalizes to unseen bodies — jon_barron · 2026-07-24
- Masked Visual Actions turns 15 hours of robot video into a zero-shot world model — jbhuang0604 · 2026-07-24
- Masked visual actions help video-action models generalize across robots and even orangutans — CSProfKGD · 2026-07-25