NVIDIA's Long-WAM scales world-action model context, hitting 95% on dynamic cup stacking
nvidia · hf · 2026-10-08
NVIDIA presents Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time robot control. Key finding: longer histories pay off only when the video foundation is pretrained autoregressively.
Highlights:
- Causal prediction learned from robot and egocentric videos without action labels, preserved during world-action adaptation;
- On RoboCasa GR-1, growing context from 0 to 19.2s raises success from 63.3% to 78.7%; bidirectionally pretrained initialization shows no net gain;
- Best results on LIBERO-Long, RoboTwin 2.0, and DOMINO;
- Streaming encoding, async execution, and hardware acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor (107.4ms per action chunk incl. future-video latent prediction on RTX 5090);
- Real-time on Unitree G1 and YAM, with 95% success on dynamic cup stacking where Pi0.5 and Fast-WAM fail all 20 trials;
- Works as a memory-informed executor alongside higher-level planning.
More from Embodied
- HKUST-GZ's UniWAM Unifies Physical Reasoning, World Generation and Action Prediction, Finds Co-Training Scaling Law — HKUSTGZ · 2026-10-08
- PKU's ViGAR Hierarchical World-Action Model Boosts Robot Manipulation Success by 12.86 Points — DAGroup-PKU · 2026-10-08
- Open-source agentic workbench: a harness that hears, speaks and sees to help assemble electronics — andreisavu · 2026-10-08
- RoboQuest Benchmark: Best Multimodal Agent Succeeds in Only 23% of Embodied Exploration Tasks — declare-lab · 2026-10-08
- RobotWorld Benchmark Tests Multimodal Agents on 84 Physical Robot Tasks — Zhiqin Yang · 2026-10-08
- Donut Robotics tests 170cm humanoid in Japanese eldercare across 100+ facilities — CyberRobooo · 2026-10-08