Masked Visual Actions turns 15 hours of robot video into a zero-shot world model

jbhuang0604 · x · 2026-07-24

Masked Visual Actions turns video models into robot world models

Researchers introduce Masked Visual Actions (MVA), a way to repurpose pre-trained video models as robot world models. The core idea is to express actions in the language of vision, letting the model answer counterfactual questions such as: if an object moves in a certain way in the clip, how will the scene evolve?

The method connects two robot-relevant capabilities:

According to the authors, the approach generalizes zero-shot to unseen embodiments and scenes despite being fine-tuned on only 15 hours of single-robot data. They also say the same representation works whether the “thing” is a robot or an object, which makes it useful for model-based planning, policy evaluation, and action inference.

Related event: Robot World Model Trained with 15 Hours of Video Achieves Zero-Shot Generalization(2 posts)→

Original post →

More from Embodied

Embodied channel →