Using Video Models for Robot Control
oier_mees · x · 2026-07-13
At RSS 2026, the author presented their latest work, Mimic-Video: Video-Action Models for Generalizable Robot Control Beyond VLAs.
The core question: What happens if a robot's policy is built on a pretrained video model rather than a static vision-language backbone?
They propose Video-Action Models (VAMs), leveraging rich temporal representations learned from large-scale video data to boost robot learning. The paper claims this approach yields better sample efficiency and training speeds compared to traditional Vision-Language-Action (VLA) models, while also benefiting from ongoing advancements in foundational video models.
The author also mentions they are in Sydney for RSS, presenting first at the Pioneers Workshop and then at the main conference, and welcomes offline chats about robot learning and foundation models.
Related event: Mimic-Video: Video-Action Robot Control Showcased at RSS 2026(2 posts)→
More from Embodied
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11