LingBot-VA 2.0 Natively Pretrained for Control

omarsar0 · x · 2026-07-11

The post explains that many video-action robot models are essentially "video generators" designed for content creation with an external action head attached. In contrast, LingBot-VA 2.0 is built from scratch, pretraining the entire video-action stack as a control-oriented system. Replies emphasize that this design is more rational than forcefully modifying existing video models into robot controllers, bringing several technical improvements.

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →