Video-Action Foundation Model for Robot Control
omarsar0 · x · 2026-07-11
Proposes a vision-action tokenizer that maps world states and latent actions into a shared semantic latent space, aligning with a frozen vision foundation model.
The core idea is that the model learns latent actions via self-supervision from unlabeled videos, reducing reliance on scarce robot demonstration data and enabling web video to provide control signals.
Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→
More from Embodied
- The Humanoid AI raises $152M Series A at a $1.35B valuation — RazRazcle · 2026-07-22
- SceniX joins World Labs to close the real-to-sim gap for robot learning — davidyin44 · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- A quadruped robot gets a custom glow-up with a new shell and screen — DynamicWebPaige · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22