NVIDIA's Long-WAM: 19.2s visual context lifts robot control from 66.3% to 78.7%

arankomatsuzaki · x · 2026-10-09

NVIDIA with MIT/HKU/UCSD releases Long-WAM, scaling the visual history of causal world-action models under real-time control constraints.

Approach: learn causal prediction from robot and egocentric videos without action labels, then preserve the history-to-future structure during world-action adaptation. Longer histories amplify the benefit of AR video pretraining; bidirectional initialization shows no net gain.

Key numbers:

Project page, code, and models are open.

Related event: NVIDIA Releases Long-WAM: 19.2-Second Visual Context Boosts Real-Time Robot Control(5 posts)→

Original post →

More from Embodied

Embodied channel →