SG-WAM: 0.9B Embodied Model Achieves 98.5% on LIBERO Benchmark
Ruiteng Zhao · hf · 2026-08-05
Introduces SG-WAM, a self-guided World Action Model framework. While existing models struggle to balance action alignment and geometric awareness when predicting future states, SG-WAM learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space.
The framework introduces learnable dynamics tokens and a Self-Guided Predictor. It uses an exponential moving average (EMA) copy of the policy backbone for stable supervision, jointly optimizing latent future prediction, geometric grounding, and flow-matching action generation end-to-end.
Experiments show that without large-scale embodied pretraining, a 0.9B parameter model achieves a 98.5% average success rate on the LIBERO benchmark and 73% on LIBERO-Plus, outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.
More from Embodied
- Tesla's Planned 'Terafab' to Span 10x Giga Texas, Target 1TW AI Compute — XFreeze · 2026-08-05
- NeurIPS 2026 Calls for Papers on Embodied Spatial Reasoning and World Models — du_yilun · 2026-08-05
- Embodied AI in Action: Agent Autonomously Plans 3D Printer Transport — neurosp1ke · 2026-08-05
- Robotics Needs Custom Small VLMs: Fine-tuned Qwen Beats APIs in Cost and Accuracy — ChongZzZhang · 2026-08-05
- Silico makes robotics model 40% more efficient by removing attention layers — mathildepapillo · 2026-08-05
- Pruning Attention Layers Slashes Action Expert FLOPs by 94% in Robotics — mathildepapillo · 2026-08-05