DreamZero: 14B Autoregressive Video Diffusion Model Enables 7Hz Real-Time Closed-Loop Robot Control
du_yilun · x · 2026-08-13
The author shares a new paper, World Action Models are Zero-shot Policies, and its code, introducing DreamZero, a World Action Model (WAM) built upon a pretrained video diffusion backbone.
- Core Strength: By jointly modeling video and action to learn physical dynamics without relying on repetitive demonstrations, DreamZero achieves over a 2x improvement in generalization to new tasks and environments compared to state-of-the-art Vision-Language-Action (VLA) models.
- System Optimization: Through model and system optimizations, the researchers enabled a 14B autoregressive video diffusion model to perform real-time closed-loop control at 7Hz.
- Cross-embodiment Transfer: It demonstrates robust cross-embodiment capabilities. Using just 10-20 minutes of video-only demonstrations from other robots or humans yields a relative improvement of over 42% on unseen tasks.
More from Embodied
- Walden Robotics Opens New SF HQ, Hiring Across Multimodal, RL and Hardware Roles — adnothing · 2026-08-14
- Uber Teams Up with Wayve and Nissan to Bring Robotaxis to Tokyo by 2026 — alexgkendall · 2026-08-14
- NVIDIA Cosmos Labs unveils WAMs and VLAs for robot learning — NVIDIAAI · 2026-08-14
- LDA-1B: A 1B Robot Foundation Model Trained on 30k Hours of Embodied Data — chris_j_paxton · 2026-08-13
- IDUN Launches Brain Sensing SDK for Earbuds with Edge Processing and Neuro-Adaptive UX — melnykowycz · 2026-08-13
- Indie Developer Attempts Token Launch to Fund Embodied Robot Evaluations — du_yilun · 2026-08-13