LingBot-VA 2.0 Embodied Model Released

JaynitMakwana · x · 2026-07-10

LingBot-VA 2.0 features a joint video-action modeling approach that predicts the next state while simultaneously planning actions, which the authors argue is more natural than traditional reactive robotic models.

Described as the first embodied-native foundation model rather than a fine-tuned video generation model, it reportedly achieves a 93.6% success rate on bimanual tasks, 150Hz inference on a single GPU, and generalization from 20 demos.

Related event: Ant Lingbo Open-Sources LingBot-VLA 2.0 for Multi-Robot Generalization(18 posts)→

Original post →

More from Embodied

Embodied channel →