Masked Visual Actions turns 15 hours of robot video into a zero-shot world model
jbhuang0604 · x · 2026-07-24
Masked Visual Actions turns video models into robot world models
Researchers introduce Masked Visual Actions (MVA), a way to repurpose pre-trained video models as robot world models. The core idea is to express actions in the language of vision, letting the model answer counterfactual questions such as: if an object moves in a certain way in the clip, how will the scene evolve?
The method connects two robot-relevant capabilities:
- Forward dynamics: predict future observations from observations and actions
- Inverse dynamics: infer actions from observation sequences
According to the authors, the approach generalizes zero-shot to unseen embodiments and scenes despite being fine-tuned on only 15 hours of single-robot data. They also say the same representation works whether the “thing” is a robot or an object, which makes it useful for model-based planning, policy evaluation, and action inference.
Related event: Zero-Shot Robot World Model Trained with 15 Hours of Video(3 posts)→
More from Embodied
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11