Simple-WAM: One forward pass replaces video denoising, keeping world action models' generalization

Tsinghua-LeapLab · hf · 2026-09-30

Tsinghua LeapLab studies whether world action models (WAMs) must explicitly generate future frames at inference. Latent WAMs, which discard the future for speed, match explicit ones in-distribution but consistently lose generalization benefits across environmental perturbation, data efficiency, and task generalization. The authors show the gap comes almost entirely from the first denoising step—the benefit comes from preparing the future, not generating it. They propose Simple-WAM, which reduces future modeling to a single forward pass over fully noised video tokens with an adapted noise schedule, achieving explicit-level generalization at latent-level efficiency across simulation and real-world tasks.

Original post →

More from Embodied

Embodied channel →