Meta's NAVA-WAM pretrains robot action policies directly from action-free videos
meta · hf · 2026-10-05
Meta researchers introduced NAVA-WAM, a world action model that learns action priors natively from observation-only videos, bypassing the need for action-annotated robot trajectories.
- Approach: Instead of pretraining visual representations for later control adaptation or inferring latent actions, the model pretrains the action policy (Action-DiT) directly, using future-video flow-matching supervision propagated through transition-structured joint attention.
- Two-stage training: Stage one pretrains on action-free videos to learn action-relevant priors; stage two post-trains with action-labeled demos via joint video-action flow matching, with asymmetric attention decoupling the visual branch for efficient action-only inference.
- Results: Consistently outperforms prior methods in both in-distribution and out-of-distribution settings, with strong action-label efficiency and real-robot generalization.
The work establishes native action-prior learning as a scalable path beyond action-labeled robot data.
More from Embodied
- Physical AI field session in Bengaluru to tackle post-deployment evals and retraining loops — carrycooldude · 2026-10-05
- Robotics' most-hyped model completed just 7 of 100 real manipulation tasks, demo reel hides the rest — carrycooldude · 2026-10-05
- Sergey Levine on robotics: hardware is good enough, the real gap is decision-making and data — 机器之心 · 2026-10-05
- SJTU's LIFT adds force sensing to VLAs with zero force-labeled pretraining data — jiqizhixin · 2026-10-05
- Rebuilding the Muse AI toy on M5Stack with Muse Gadgets SDK, customized to Chinese UI — op7418 · 2026-10-05
- Guangzhou hotel deploys wheeled humanoid robots to greet guests and brew coffee — CyberRobooo · 2026-10-05