MaskedVisualActions turns robot motion into pixels and lifts RoboCasa success by up to 26 points
机器之心 · wechat · 2026-07-23
A Stanford, Maryland, and Harvard team proposes MaskedVisualActions, a way to render robot actions as pixel-level motion trails so a video world model can reason about them directly.
What it does
- Converts robot candidate actions into visual trajectories instead of numeric control states.
- Lets the same model do forward prediction (action future video) and inverse generation (desired object motion possible robot behavior).
- Supports different robot embodiments with the same visual interface.
Results
- Fine-tuned with about 15 hours of robot interaction data.
- In RoboCasa, video-model-assisted planning improved success rates by 7 to 26 percentage points.
- Predicted strategy performance correlated with real simulation at 0.982.
- The inverse-action pipeline reached 90% success on CoffeeServeMug.
Caveats
The model is still limited by the underlying video model: it can be optimistic, produce physically inconsistent futures, and needs extra inverse-dynamics machinery for execution.
Related event: MaskedVisualActions Turns Robot Actions into Pixel Trajectories(2 posts)→
More from Embodied
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11