4D-WAM: Infusing Spatiotemporal Awareness into World Action Models via Trajectory Fields

zhenjun_zhao · x · 2026-08-13

World-Action Models (WAMs) typically represent videos in 2D pixel space, creating a representation gap with the 3D space where robotic actions are executed. Existing 3D approaches fail to fully exploit the dynamics of 3D structures. This paper proposes 4D-WAM, a model-agnostic training strategy.

Original post →

More from Embodied

Embodied channel →