Video-Action Foundation Model for Robot Control

omarsar0 · x · 2026-07-11

Proposes a vision-action tokenizer that maps world states and latent actions into a shared semantic latent space, aligning with a frozen vision foundation model.

The core idea is that the model learns latent actions via self-supervision from unlabeled videos, reducing reliance on scarce robot demonstration data and enabling web video to provide control signals.

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →