Faster World Action Models for Robotics
open-gigaai · hf · 2026-07-16
This work presents improvements to the **World Action Model (WAM)** in robot policy learning. The original WAM explicitly generates future videos during inference, which aids action learning but incurs massive overhead, hindering real-time closed-loop control. The authors propose **GigaWorld-Policy-0.5**, decoupling training and inference: - **Training Phase**: Combines Action-Conditioned World Modeling with WAM training to tightly couple visual dynamics with action representations - **Inference Phase**: Performs action decoding only, eliminating future video generation to reduce overhead - Introduces **Mixture-of-Transformers** to separate visual dynamics modeling and action generation into specialized experts, boosting inference efficiency - Employs an **agent-based AutoResearch** pipeline to automatically search training configurations, minimizing manual tuning The authors report that this method achieves an action inference latency of **85ms** on a local **RTX 4090**. Experiments and ablations demonstrate enhanced real-time robot control while retaining the training benefits of future visual dynamics.
More from Embodied
- MW team shows Gen1 of MW-bot, a semi-humanoid home robot built for pantry storage and ceiling rails — CyberRobooo · 2026-07-21
- Kunlun Spins Out Riemann Dynamics, Unveils World Action Model for Robotics — 0xAllen_ · 2026-07-21
- BrainCo demos near-real-time bionics without implants and claims 85% lower prosthetic cost — TrueOrange9944 · 2026-07-21
- OpenAI’s $230 CodexMicro sold out, and users are already cloning it with Stream Decks — APPSO · 2026-07-21
- Xiaomi-Robotics-1 shows robot motion improves more from data than bigger models — The Decoder · 2026-07-21
- HarmoHOI generates multi-view hand-object videos and aligned 3D motion in one diffusion model — cn-scut · 2026-07-21