LingBot-VLA 2.0: Built for Robot Control

bendee983 · x · 2026-07-11

To address the high latency and lack of action cognition when directly adapting video models for robotics, LingBot-VLA 2.0 is trained from scratch specifically for robot control.

Its core architecture features a visual tokenizer, action prediction, and a full causal Transformer. The model leverages web-scale video (including human demonstrations) for scalable self-supervised learning to capture action-relevant visual dynamics.

Evaluation Performance:

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →