Ant Lingbo Open-Sources LingBot-VLA 2.0 for Multi-Robot Generalization

Ant Group's embodied intelligence subsidiary, Ant Lingbo, officially released and open-sourced LingBot-VLA 2.0, a Vision-Language-Action (VLA) foundation model. Tackling the core challenge of cross-configuration generalization in embodied AI, the model enables a single brain to drive multiple robot embodiments, significantly enhancing practical capabilities in real-world physical tasks.

Key Technologies and Data Scale

The LingBot-VLA 2.0 model has a scale of 6B parameters and is released under the Apache 2.0 license. Its pre-training data reaches approximately 60,000 hours, comprising 50,000 hours of real robot trajectories and 10,000 hours of egocentric human operation videos. Technically, the model utilizes a sparse MoE design and token-level MoE action head, combined with multi-teacher distillation, achieving higher training efficiency and lower errors under the same activated parameter budget. Additionally, it introduces causal world modeling and future prediction capabilities, enabling simultaneous next-step prediction and action planning through joint video-action modeling.

Cross-Configuration Generalization and Performance

The model expands the action space to full-body degrees of freedom, including the head, waist, end effectors, and mobile base. It successfully unifies 20 different robot configurations across 17 mainstream domestic and international brands (such as Unitree, AGIBOT, and Astribot) into a single action space. In tests, a single training policy autonomously drove various robots ranging from the Franka single-arm to Fourier GR-2 and Unitree G1 humanoids. It demonstrated superior success rates and task progress in bimanual collaborative tasks and long-horizon mobile operations compared to previous generations.

2026-07-08 ~ 2026-07-10 · 18 related posts

1 near-duplicate retellings: 新智元