LingBot-VA 2.0 Released

chris_j_paxton · x · 2026-07-10

LingBot-VA 2.0 has been released as a native video-action foundation model designed for general robot control. The author emphasizes that rather than repurposing a general video generation model, it was trained from scratch using joint video-action pre-training.

The post outlines three core aspects: native video-action pre-training, a semantic visual-action tokenizer, and foresight reasoning for advance action planning. The model can run in real-time on consumer-grade GPUs, supports up to 150Hz control, and generalizes to unseen tasks.

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →