BAAI's World Action Models use video generation to let robots imagine before acting

stepjamUK · x · 2026-09-15

A team from BAAI and collaborators proposes World Action Models (WAM): video generation models that let a robot imagine what happens next before deciding how to act, tackling the tension that prediction takes time a moving robot doesn't have.

Key insight on latency vs. fidelity: each denoising round adds latency to the control loop, so the obvious fix is one round. But watching the denoising process, the team found the scene background sharpens almost immediately while the gripper, the object, and their interaction stay blurry until several steps later. Cutting the process short keeps a crisp picture of the room and loses exactly the details that matter for manipulation — explaining why naive denoising reduction quietly breaks robot control.

Original post →

More from Embodied

Embodied channel →