World Action Models survey says robotics is moving from reacting to predicting consequences

rohanpaul_ai · x · 2026-07-27

This survey on World Action Models (WAMs) says robotics is shifting from reacting to the present toward predicting the consequences of action before acting.

The paper defines a WAM as a model whose predicted future directly helps produce, score, verify, or train the action. It argues that “dream less, act more”: full video generation is often too slow and memory-heavy for control loops, so newer systems increasingly use latent features, geometry, affordance maps, motion representations, tactile signals, and other physically grounded representations instead of rendered video.

The survey also lays out the main trade-offs—predictive richness versus latency, memory, action-label cost, and reliability—and says there is no single winning architecture yet. The open question is when robots should spend heavy predictive compute only when uncertainty, contact, or irreversible error makes it necessary.

Related event: New Robotics Paradigm: Shifting from Reactive Control to Predictive Action(2 posts)→

Original post →

More from Embodied

Embodied channel →