SG-WAM: Enhancing World-Action Models with VLM-Based Semantic Guidance
zhenjun_zhao · x · 2026-08-13
Existing World-Action Models (WAMs) rely primarily on visual cues rather than language instructions, causing predicted videos to be semantically misaligned with instructions and degrading action prediction accuracy. This paper proposes SG-WAM to overcome this limitation.
- VLM Planner: Introduces a vision-language model (VLM) as a semantic planner to enhance the instruction-grounding capacity of WAMs.
- Semantic Foresight: Trains the planner to predict text-grounded and spatial-aware semantic foresight. Text-grounded foresight identifies correct target objects, while spatial-aware foresight provides scene geometry for precise manipulation.
- Injection: Injects this foresight into the WAM as high-level semantic guidance, ensuring both future-video generation and action prediction faithfully follow language instructions. Superiority is demonstrated in both simulation and real-world experiments.
More from Embodied
- VLAIRobotics Launches Dual-Arm Humanoid K1 Starting at $2,700 — AI寒武纪 · 2026-08-13
- Dyna-2 Pre-Trained on 1M Hours of Human Video Achieves Cross-Embodiment Transfer — iamfakhrealam · 2026-08-13
- Denmark Develops Jointless Earthworm-Inspired Soft Robot for Search and Rescue — lukas_m_ziegler · 2026-08-13
- SHAPER: Train-Free Skill Evolution for Embodied Agents — Peidong Wang · 2026-08-13
- Expert: Humanoid Robots Could Be Trillion-Dollar Market, But Form Factor Depends on Environment — MarwaEldiwiny · 2026-08-13
- Minimax H3 Runs on RTX 5060 Ti: 25-Second Video Takes 93 Minutes to Generate — Beginning_Tip300 · 2026-08-13