NVIDIA and MIT Release Long-WAM: 19.2s Visual Context Boosts Robot Success to 78.7%
songhan_mit · x · 2026-10-09
Researchers from NVIDIA, MIT, HKU, and UCSD released Long-WAM, a framework for scaling the visual context of causal world-action models under real-time control constraints. Highlights:
- Scaling context from 2.4s to 19.2s lifts RoboCasa GR-1 success from 66.3% to 78.7%; bidirectional initialization shows no net gain
- Autoregressive video pretraining (no action labels) is key — longer histories amplify its advantage
- Benchmarks: 99.5% LIBERO-Long, 94.4% RoboTwin 2.0, 95% dynamic cup stacking on Unitree G1 (19/20)
- Co-designed runtime runs the full predict-then-act model at 107.4 ms per action chunk on RTX 5090; deploys on DGX Spark and Jetson AGX Thor
Paper, code, and models are public on arXiv and Hugging Face.
More from Embodied
- IRVL Researcher to Present UHAS, HRT1 and VLA-Replica at Texas A&M Robotics Seminar — YuXiang_IRVL · 2026-10-09
- Google Gemma-Backed Hardware Hackathon at AGI House Gathers 120 Builders for On-Device AI — agihouse_org · 2026-10-09
- Odyssey launches Odyssey-3, claims SOTA world model on Physics-IQ benchmark — Scobleizer · 2026-10-09
- Figure valued at $39B: tracking Brett Adcock's promise of 100,000 robots in 4 years — annatonger · 2026-10-09
- Sentdex: a raw multimodal LLM drives a robot with zero training, VLA era may be over — Sentdex · 2026-10-09
- Dot users complain the iPhone app still lacks CallKit for calls — cheezemink · 2026-10-09