Is VLA Dead? World Models Achieve 1.55x Success Rate Over VLA in Complex Tasks

CyberRobooo · x · 2026-08-11

The post discusses whether VLA (Vision-Language-Action) models are being outperformed in robotics. In a head-to-head comparison by Dyna, the World Action Model (WAM) achieved 1.55 times the average success rate of VLA across 7 real-world tasks using identical pretraining data.

WAM demonstrated significant advantages in complex manipulations like bottle opening and key turning. Key turning was nearly unsolved at 100K hours of pretraining but reached a 90% success rate at 1M hours. Meanwhile, NVIDIA's EgoScale approach improved dexterous manipulation success rates by 54% by incorporating 20,000+ hours of human egocentric data.

The two approaches are converging: VLA connects human data directly to robot tasks, while WAM learns how the physical world evolves (via video prediction) to output actions.

Related event: Dyna Robotics Launches Dyna-2 Trained on 1M Hours of Human Video(27 posts)→

Original post →

More from Embodied

Embodied channel →