A 15-hour video fine-tune turns masked trajectories into a robot control interface
Hadi Alzayer · hf · 2026-07-22
Masked Visual Actions proposes a pixel-space control interface for video models that is grounded in physical manipulation.
Instead of using abstract action tokens, the method encodes action as a partially revealed trajectory of an entity in video. Revealing robot motion turns the model into a forward dynamics predictor; revealing desired object motion lets it infer robot behavior that would produce that outcome. Fine-tuned on only 15 hours of masked examples from real and simulated videos, one checkpoint shows strong fidelity and controllability across scenes and embodiments. The model can evaluate imagined rollouts against real execution, rank candidate futures for planning, and synthesize robot motion from desired object motion.
Related event: MaskedVisualActions Turns Robot Actions into Pixel Trajectories(2 posts)→
More from Embodied
- Amazon and Google sold 600M+ smart speakers, so why no AGI-era successor? — julianlehr · 2026-09-11
- ECCV26 Oral: Flow Matching Enables Single-Stage Multi-View Point Cloud Registration — ducha_aiki · 2026-09-11
- Polish developers build iPhone app that detects nearby Meta smart glasses — Low-Honeydew6483 · 2026-09-11
- Ant's Afu health AI hits 150M users, unveils AI+hardware health alliance at Bund Summit — APPSO · 2026-09-11
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11