Masked Visual Actions: Turning Video Models into Robotic World Models
jiqizhixin · x · 2026-08-03
Researchers from Stanford, University of Maryland, and Harvard introduced Masked Visual Actions (MVA), a novel method to transform video models into robotic world models.
By revealing partial robot motion in pixel space (masking parts of a video), the model learns to predict how scenes respond to actions and determine the necessary actions to achieve a goal.
Key Highlights:
- Data Efficient: Trained on just 15 hours of data.
- Unified Capabilities: A single checkpoint handles policy evaluation, model-based planning, and action extraction.
- Strong Generalization: Capable of zero-shot generalization across diverse environments and multiple robot embodiments.
More from Embodied
- ABB and Massive Dimension Unveil 6-Axis 3D Printing Robot — lukas_m_ziegler · 2026-08-03
- Open Source Tool Lichtblick: In-Browser Robotics Data Visualization and Diagnostics — tom_doerr · 2026-08-03
- Y Combinator Showcases Embodied Robot Cosmic-1 — ycombinator · 2026-08-03
- Robot Weaponization Regulation Lags Behind LLM Safety Standards — alex_verem · 2026-08-03
- Robot Deployment Pain Points: Choosing the Right 3D Printing Material for Grippers — DominiqueCAPaul · 2026-08-03
- Lambda Surgical Raises $6.5M to Build World Models for the Human Body — ditzikow · 2026-08-03