Using Video Models for Robot Control
oier_mees · x · 2026-07-13
At RSS 2026, the author presented their latest work, Mimic-Video: Video-Action Models for Generalizable Robot Control Beyond VLAs.
The core question: What happens if a robot's policy is built on a pretrained video model rather than a static vision-language backbone?
They propose Video-Action Models (VAMs), leveraging rich temporal representations learned from large-scale video data to boost robot learning. The paper claims this approach yields better sample efficiency and training speeds compared to traditional Vision-Language-Action (VLA) models, while also benefiting from ongoing advancements in foundational video models.
The author also mentions they are in Sydney for RSS, presenting first at the Pioneers Workshop and then at the main conference, and welcomes offline chats about robot learning and foundation models.
Related event: Mimic-Video: Video-Action Robot Control Showcased at RSS 2026(2 posts)→
More from Embodied
- Polymarket puts Tesla’s California robotaxi launch odds at 16% this year — Polymarket · 2026-07-21
- Tesla expands robotaxi service to Orlando and Tampa — Polymarket · 2026-07-21
- Humanoid robots are approaching a deeper uncanny valley — GlenBradley · 2026-07-21
- Polymarket gives Tesla’s Optimus just a 17% chance of debuting this year — Polymarket · 2026-07-21
- UK robotics startup Humanoid raises $152 million at a $1.35 billion valuation — Polymarket · 2026-07-21
- Halliday opens priority access to G2 display AI glasses for meetings and daily use — SucceededMind · 2026-07-21