BAAI's World Action Models use video generation to let robots imagine before acting
stepjamUK · x · 2026-09-15
A team from BAAI and collaborators proposes World Action Models (WAM): video generation models that let a robot imagine what happens next before deciding how to act, tackling the tension that prediction takes time a moving robot doesn't have.
Key insight on latency vs. fidelity: each denoising round adds latency to the control loop, so the obvious fix is one round. But watching the denoising process, the team found the scene background sharpens almost immediately while the gripper, the object, and their interaction stay blurry until several steps later. Cutting the process short keeps a crisp picture of the room and loses exactly the details that matter for manipulation — explaining why naive denoising reduction quietly breaks robot control.
More from Embodied
- Verne Robotics bets on on-device inference: robots can't rely on WiFi to run policies — danfei_xu · 2026-09-15
- Robot Foundation Models Like Astra May Be Key to Truly Safe Robots — chris_j_paxton · 2026-09-15
- Moonphase Flip: a low-distraction AI to-do screen on the back of your phone — future_coded · 2026-09-15
- AI model paints Golden Gate Bridge with a real robot arm, now reproduced in MuJoCo and open-sourced — victormustar · 2026-09-15
- Reimagine Robotics builds for a future where customers teach robots on the job — notmisha · 2026-09-15
- SimpliSafe's new $199.99 AI doorbell pairs on-device detection with live human guards — The Verge AI · 2026-09-15