4D-WAM: Infusing Spatiotemporal Awareness into World Action Models via Trajectory Fields
zhenjun_zhao · x · 2026-08-13
World-Action Models (WAMs) typically represent videos in 2D pixel space, creating a representation gap with the 3D space where robotic actions are executed. Existing 3D approaches fail to fully exploit the dynamics of 3D structures. This paper proposes 4D-WAM, a model-agnostic training strategy.
- Core Method: Injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment.
- Dual Objectives: 1) Motion alignment aligns temporal feature variations across adjacent frames to build local 4D awareness; 2) Destination alignment guides the model to infer the final destination by minimizing the gap between attention-like similarity distributions of the source and destination frames.
- Results: These objectives provide local motion supervision and long-horizon goal guidance, enabling WAMs to learn trajectory-level spatiotemporal representations, performing well in in-distribution and out-of-distribution experiments.
More from Embodied
- VLAIRobotics Launches Dual-Arm Humanoid K1 Starting at $2,700 — AI寒武纪 · 2026-08-13
- Dyna-2 Pre-Trained on 1M Hours of Human Video Achieves Cross-Embodiment Transfer — iamfakhrealam · 2026-08-13
- Denmark Develops Jointless Earthworm-Inspired Soft Robot for Search and Rescue — lukas_m_ziegler · 2026-08-13
- SHAPER: Train-Free Skill Evolution for Embodied Agents — Peidong Wang · 2026-08-13
- Expert: Humanoid Robots Could Be Trillion-Dollar Market, But Form Factor Depends on Environment — MarwaEldiwiny · 2026-08-13
- Minimax H3 Runs on RTX 5060 Ti: 25-Second Video Takes 93 Minutes to Generate — Beginning_Tip300 · 2026-08-13