Using Video Models for Robot Control

oier_mees · x · 2026-07-13

At RSS 2026, the author presented their latest work, Mimic-Video: Video-Action Models for Generalizable Robot Control Beyond VLAs.

The core question: What happens if a robot's policy is built on a pretrained video model rather than a static vision-language backbone?

They propose Video-Action Models (VAMs), leveraging rich temporal representations learned from large-scale video data to boost robot learning. The paper claims this approach yields better sample efficiency and training speeds compared to traditional Vision-Language-Action (VLA) models, while also benefiting from ongoing advancements in foundational video models.

The author also mentions they are in Sydney for RSS, presenting first at the Pioneers Workshop and then at the main conference, and welcomes offline chats about robot learning and foundation models.

Related event: Mimic-Video: Video-Action Robot Control Showcased at RSS 2026(2 posts)→

Original post →

More from Embodied

Embodied channel →