DeltaWAM cuts video model cost for bimanual manipulation, lifting RoboTwin success to 85.4%
Han Yan · hf · 2026-09-25
Researchers propose DeltaWAM, a Delta World Action Model for bimanual manipulation that fixes existing world-action models' wasteful dense future-frame prediction.
- Method: jointly predicts visual deltas and actions via dense-anchor, sparse-delta, and action streams, with three architecture variants differing in representation and computation sharing.
- Streaming Delta Memory (SDM): updates cached anchor context with compact observed deltas, avoiding heavy video-expert processing each step.
- Results: on RoboTwin, average success over Fast-WAM improves from 81.3% to 85.4% clean and from 75.8% to 83.9% under visual randomization; the three architectures cut training FLOPs by 17.78–23.77%, SDM reduces one-step inference latency and FLOPs by 36.57% and 31.55%; real-world evaluations show the highest overall success and normalized progress.
Code and project page are open-sourced.
More from Embodied
- Robots & Beers meetup returns to Hong Kong Oct 15 as an Electronics Fair side event — broodsugar · 2026-09-25
- XPENG locks in suppliers for humanoid robots, targets 2027 deliveries — VraserX · 2026-09-25
- Sharpa's World Synesthesia Model for dexterous hands accepted at CoRL 2026 — jiqizhixin · 2026-09-25
- IFR report: 5M robots now working, China alone accounts for 59% of 2025 installs — lukas_m_ziegler · 2026-09-25
- Musk on China's humanoid robot games: the future will have 'a lot, a lot' of robots — XFreeze · 2026-09-25
- Meta sidesteps the VR/XR/MR jargon mess by just calling them VR glasses — pvncher · 2026-09-25