MaskedVisualActions turns robot motion into pixels and lifts RoboCasa success by up to 26 points

机器之心 · wechat · 2026-07-23

A Stanford, Maryland, and Harvard team proposes MaskedVisualActions, a way to render robot actions as pixel-level motion trails so a video world model can reason about them directly.

What it does

Results

Caveats

The model is still limited by the underlying video model: it can be optimistic, produce physically inconsistent futures, and needs extra inverse-dynamics machinery for execution.

Related event: MaskedVisualActions Turns Robot Actions into Pixel Trajectories(2 posts)→

Original post →

More from Embodied

Embodied channel →