Robot World Model Trained with 15 Hours of Video Achieves Zero-Shot Generalization

Researchers introduced Masked Visual Actions (MVA), transforming pre-trained video models into robot world models. Trained on just 15 hours of single-arm robot video, the model achieves zero-shot generalization to unseen embodiments, including orangutans.

2026-07-24 ~ 2026-07-24 · 2 related posts