Masked Visual Actions turns 15 hours of robot video into a zero-shot world model
jbhuang0604 · x · 2026-07-24
Masked Visual Actions turns video models into robot world models
Researchers introduce Masked Visual Actions (MVA), a way to repurpose pre-trained video models as robot world models. The core idea is to express actions in the language of vision, letting the model answer counterfactual questions such as: if an object moves in a certain way in the clip, how will the scene evolve?
The method connects two robot-relevant capabilities:
- Forward dynamics: predict future observations from observations and actions
- Inverse dynamics: infer actions from observation sequences
According to the authors, the approach generalizes zero-shot to unseen embodiments and scenes despite being fine-tuned on only 15 hours of single-robot data. They also say the same representation works whether the “thing” is a robot or an object, which makes it useful for model-based planning, policy evaluation, and action inference.
More from Embodied
- Software Moats are Gone: Founders Bet on Hardware + AI Strategy — pramodk73 · 2026-07-24
- Bifrost says its real2sim pipeline can rebuild a site video into a simulation-ready 3D world in 30 minutes — jnack · 2026-07-24
- A retro wrist computer becomes a joke about humanoid robot emergency overrides — seanmcdonaldxyz · 2026-07-24
- Vision-language-action systems can still miss the next action chunk — StillThese3747 · 2026-07-24
- Shengshu unveils Vidu, ViduS1 and Motubrain as a full world-model stack at WAIC 2026 — 生数科技 · 2026-07-24
- OpenAI could become the intelligence layer for dozens of robot brands — VraserX · 2026-07-24