EVO-WAM self-verifies video-action rollouts, real-world long-horizon success 20%→76.7%
hitsz123 · hf · 2026-09-30
EVO-WAM adapts world action models (WAMs) to unseen tasks using their own generated video-action trajectories—no extra expert demos or external execution needed.
- Adds state prediction and anchored multi-frame context for full autoregressive rollouts; a VLM selects task-completing prefixes and an inverse dynamics model verifies video-action consistency before training; iterates training and rollout generation
- On seven unseen RoboTwin 2.0 tasks, average success rises from 26.9% to 68.0% for Cosmos3 (2.5x) and 28.5% to 46.4% for DreamZero (1.6x)
- On three real-world long-horizon composite tasks, Cosmos3 improves from 20.0% to 76.7%, a 56.7-point gain
More from Embodied
- DIY open-source driving mods: Tesla owner runs Sunnypilot on Model Y with $1,000 Comma Four kit, alarming experts — science · 2026-09-30
- Call for physical AI founders: actuators, dexterous hands, robot foundation models and more — Sethwinterroth · 2026-09-30
- WorldLine: 10k-hour video-trained action simulator lifts robot policy success up to 21.4 pts — hongkongust · 2026-09-30
- PanoVLN beats VLN SOTA by 11.9% success rate using panoramic views, tested on a quadruped robot — ZhejiangUniversity · 2026-09-30
- Agile Robots ships humanoid robots daily from German factory, partners with DeepMind — CyberRobooo · 2026-09-30
- Real2Gym Turns Videos Into Interactive Robot Gyms, Beating GPT-6 Direct Mode by 33% on Real Robots — Shanghai-AI-Laboratory · 2026-09-30