VioLA trains humanoid policies on 140M mostly-human frames, hitting 88.6% zero-shot manipulation
Michael_J_Black · x · 2026-10-10
Researchers from Vesoma AI, MPI-IS, ETH and others released VioLA: Learning Generalist Humanoid Control Policies from Human Data on arXiv.
- Key idea: humanoid joint action spaces are large and tightly coupled, and robot demos are scarce. VioLA predicts body- and hand-motion latents instead of joint commands, with pretrained controllers executing them; human and robot motions are encoded into the same latent spaces, so human recordings directly supervise the policy.
- Scale: the training pool contains 140.6M frames, 93.2% of them human.
- Results: 100% zero-shot success on real-robot locomotion instructions (vs 16.7% for GR00T N1.7 and 0% for Ψ₀), and 88.6% manipulation success without task-specific fine-tuning.
Authors include Michael J. Black, Andreas Krause, Georg Martius and Martin Riedmiller.
More from Embodied
- FIND preprint: robot picks its own weaknesses, lifts 8-task success from 55% to 71.9% — GeorgiaChal · 2026-10-10
- Musk: Digital Optimus beats Diablo halfway through with no APIs, just screen pixels — XFreeze · 2026-10-10
- Tongji team builds spiderweb-inspired 6-axis robotic 3D printer with self-supporting structures — lukas_m_ziegler · 2026-10-10
- Robot hand turns book pages: UVTA visual-tactile-action model hits 70% vs 29% baseline — CyberRobooo · 2026-10-10
- Sony surgical robot's autonomous suturing demo sparks debate over medical AI liability — TansuYegen · 2026-10-10
- UPenn's GRASP lab releases OctoSense: 8.5TB multi-sensor robotics dataset with event cameras — RexDouglass · 2026-10-10