SG-WAM: 0.9B Embodied Model Achieves 98.5% on LIBERO Benchmark

Ruiteng Zhao · hf · 2026-08-05

Introduces SG-WAM, a self-guided World Action Model framework. While existing models struggle to balance action alignment and geometric awareness when predicting future states, SG-WAM learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space.

The framework introduces learnable dynamics tokens and a Self-Guided Predictor. It uses an exponential moving average (EMA) copy of the policy backbone for stable supervision, jointly optimizing latent future prediction, geometric grounding, and flow-matching action generation end-to-end.

Experiments show that without large-scale embodied pretraining, a 0.9B parameter model achieves a 98.5% average success rate on the LIBERO benchmark and 73% on LIBERO-Plus, outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.

Original post →

More from Embodied

Embodied channel →