NVIDIA's Long-WAM: 19.2s visual context lifts robot control from 66.3% to 78.7%
arankomatsuzaki · x · 2026-10-09
NVIDIA with MIT/HKU/UCSD releases Long-WAM, scaling the visual history of causal world-action models under real-time control constraints.
Approach: learn causal prediction from robot and egocentric videos without action labels, then preserve the history-to-future structure during world-action adaptation. Longer histories amplify the benefit of AR video pretraining; bidirectional initialization shows no net gain.
Key numbers:
- RoboCasa GR-1: context from 2.4s to 19.2s raises success from 66.3% to 78.7%
- 99.5% on LIBERO-Long, 94.4% average on RoboTwin 2.0
- Co-designed runtime runs the full predict-then-act model at 107.4ms per action chunk on an RTX 5090, deployable on DGX Spark and Jetson AGX Thor
- 95% dynamic cup stacking on Unitree G1 (19/20); 54.4% on RoboCasa365 with GPT-6 Astra across 50 tasks
Project page, code, and models are open.
More from Embodied
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- Sentdex: a raw multimodal LLM drives a robot with zero training, VLA era may be over — Sentdex · 2026-10-09
- Figure's $39B bet: how far has Brett Adcock gotten on his 100,000-robot promise? — annatonger · 2026-10-09
- Dot users complain the iPhone app still lacks CallKit for calls — cheezemink · 2026-10-09
- BiGym 2.0: LLM-Written Policies Hit 65% From One Demo, But π0.5 Still Leads at 75% — stepjamUK · 2026-10-09
- Dev builds interactive teardown of a humanoid robot with 56 actuated joints using Claude Opus — techartist_ · 2026-10-09