Ant Lingbo Open-Sources LingBot-VLA 2.0: General Robot Brain

量子位 · wechat · 2026-07-08

Ant Lingbo has released and open-sourced LingBot-VLA 2.0, a general Vision-Language-Action (VLA) model designed for complex physical world tasks. Pre-trained on 60000 hours of data (50000 hours of real robot trajectories + 10000 hours of first-person human operation videos), it supports 20 robot configurations. Its action space has been expanded from dual arms to include the head, waist, mobile base, and dexterous hands.

Integrating the LingBot-Depth spatial understanding module, the model achieves an inference latency of less than 130ms on an RTX 4090. It also introduces a future action prediction mechanism to improve coherence in long-sequence tasks. Demo videos show the robot smoothly completing long-sequence household chores such as organizing the fridge, cleaning the stovetop, and sorting condiments.

Compared to the first generation, the data scale has expanded from 20000 to 60000 hours in just six months, covering more robot forms and task types. This demonstrates the continued progress of the "real data scale driving general manipulation capabilities" approach.

Related event: Ant Lingbo Open-Sources LingBot-VLA 2.0 for Multi-Robot Generalization(18 posts)→

Original post →

More from Embodied

Embodied channel →