Simple-WAM: One forward pass replaces video denoising, keeping world action models' generalization
Tsinghua-LeapLab · hf · 2026-09-30
Tsinghua LeapLab studies whether world action models (WAMs) must explicitly generate future frames at inference. Latent WAMs, which discard the future for speed, match explicit ones in-distribution but consistently lose generalization benefits across environmental perturbation, data efficiency, and task generalization. The authors show the gap comes almost entirely from the first denoising step—the benefit comes from preparing the future, not generating it. They propose Simple-WAM, which reduces future modeling to a single forward pass over fully noised video tokens with an adapted noise schedule, achieving explicit-level generalization at latent-level efficiency across simulation and real-world tasks.
More from Embodied
- USC's CLAM learns robot policies from unlabeled videos, 2-3x success over baselines — ebiyik_ · 2026-09-30
- NVIDIA's Robo Olympics uses Codex and natural language to train robot skills in physics simulation — RevLebaredian · 2026-09-30
- A huggable humanoid robot called Baymax shows up at IROS 2026 — CyberRobooo · 2026-09-30
- VLANeXt Family: 500+ controlled experiments distill 12 practical recipes for VLA models — ccloy · 2026-09-30
- Open-source humanoid Roboparty shows up at IROS, alongside a Gundam you can touch — 4310sy · 2026-09-30
- Scoble says Google Glass is coming back: 'We were early, not wrong' — Scobleizer · 2026-09-30