MaskedVisualActions turns robot motion into pixels and lifts RoboCasa success by up to 26 points
机器之心 · wechat · 2026-07-23
A Stanford, Maryland, and Harvard team proposes MaskedVisualActions, a way to render robot actions as pixel-level motion trails so a video world model can reason about them directly.
What it does
- Converts robot candidate actions into visual trajectories instead of numeric control states.
- Lets the same model do forward prediction (action future video) and inverse generation (desired object motion possible robot behavior).
- Supports different robot embodiments with the same visual interface.
Results
- Fine-tuned with about 15 hours of robot interaction data.
- In RoboCasa, video-model-assisted planning improved success rates by 7 to 26 percentage points.
- Predicted strategy performance correlated with real simulation at 0.982.
- The inverse-action pipeline reached 90% success on CoffeeServeMug.
Caveats
The model is still limited by the underlying video model: it can be optimistic, produce physically inconsistent futures, and needs extra inverse-dynamics machinery for execution.
Related event: MaskedVisualActions Turns Robot Actions into Pixel Trajectories(2 posts)→
More from Embodied
- Robot goes to the fridge and fetches a beer — Darpinian · 2026-07-27
- Researchers show digital circuits can be replicated with knitted fabric — mtizard · 2026-07-27
- Local Qwen models power a robot that tests 78 smartphones’ battery life — gappyvalley · 2026-07-27
- TechCrunch says brain-wave signals could be the next unlock for physical AI training — TechCrunch AI · 2026-07-27
- YC talk says robots need better pretraining, memory, and compositional behaviors — ycombinator · 2026-07-27
- China’s drone-license boom cools as training centers shrink and professional demand takes over — 创业邦 · 2026-07-27