Masked Visual Actions: Turning Video Models into Robotic World Models

jiqizhixin · x · 2026-08-03

Researchers from Stanford, University of Maryland, and Harvard introduced Masked Visual Actions (MVA), a novel method to transform video models into robotic world models.

By revealing partial robot motion in pixel space (masking parts of a video), the model learns to predict how scenes respond to actions and determine the necessary actions to achieve a goal.

Key Highlights:

Original post →

More from Embodied

Embodied channel →