Faster World Action Models for Robotics

open-gigaai · hf · 2026-07-16

This work presents improvements to the World Action Model (WAM) in robot policy learning. The original WAM explicitly generates future videos during inference, which aids action learning but incurs massive overhead, hindering real-time closed-loop control.

The authors propose GigaWorld-Policy-0.5, decoupling training and inference:

The authors report that this method achieves an action inference latency of 85ms on a local RTX 4090. Experiments and ablations demonstrate enhanced real-time robot control while retaining the training benefits of future visual dynamics.

Original post →

More from Embodied

Embodied channel →