NowWAM Swaps Future Prediction for Denoising in Robot Control, Hits 87.7% on LIBERO-Plus
Zanyi Wang · hf · 2026-10-06
NowWAM is a future-target-free co-training formulation that adapts pretrained generative DiTs to robot control by denoising the current observation and predicting actions from the same visual stream.
- Controlled comparisons show past and future visual targets perform comparably, while training only at the clean endpoint sharply reduces robustness — the denoising trajectory itself is the effective interface for control
- On LIBERO-Plus with FLUX2-Klein it reaches 87.7%, beating the future-target co-training baseline by 6.1 points while halving visual tokens (784→392) and cutting step time from 2.85s to 1.63s (1.8x speedup)
- With a pure text-to-image Z-Image backbone it still hits 87.8%, showing strong control adaptation doesn't require video or image-editing backbones
More from Embodied
- Robotics luminary Sergey Levine to headline South Park Commons fireside chat — adityaag · 2026-10-06
- Capstan Drive Prototype Shows Zero Slippage Ahead of Gimbal Motor Test — IanPritchard · 2026-10-06
- Anybotics launches Shift, an ops engine already running across 200+ robot deployments — lukas_m_ziegler · 2026-10-06
- Figure trained its F.02 robots to jump into molten steel to retire the whole fleet — lukas_m_ziegler · 2026-10-06
- Boston Dynamics gives Atlas a 13-DoF tactile hand; Runway open-sources Praxis-1 world action model — lukas_m_ziegler · 2026-10-06
- Dyna 2.1 semi-humanoid robot runs full commercial laundry shifts autonomously — lukas_m_ziegler · 2026-10-06