Ant Lingbo Open-Sources LingBot-VLA 2.0: General Robot Brain
量子位 · wechat · 2026-07-08
Ant Lingbo has released and open-sourced LingBot-VLA 2.0, a general Vision-Language-Action (VLA) model designed for complex physical world tasks. Pre-trained on 60000 hours of data (50000 hours of real robot trajectories + 10000 hours of first-person human operation videos), it supports 20 robot configurations. Its action space has been expanded from dual arms to include the head, waist, mobile base, and dexterous hands.
Integrating the LingBot-Depth spatial understanding module, the model achieves an inference latency of less than 130ms on an RTX 4090. It also introduces a future action prediction mechanism to improve coherence in long-sequence tasks. Demo videos show the robot smoothly completing long-sequence household chores such as organizing the fridge, cleaning the stovetop, and sorting condiments.
Compared to the first generation, the data scale has expanded from 20000 to 60000 hours in just six months, covering more robot forms and task types. This demonstrates the continued progress of the "real data scale driving general manipulation capabilities" approach.
Related event: Ant Lingbo Open-Sources LingBot-VLA 2.0 for Multi-Robot Generalization(18 posts)→
More from Embodied
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- A quadruped robot gets a custom glow-up with a new shell and screen — DynamicWebPaige · 2026-07-22
- A VR teleop demo for an SO-101 arm gets absurdly low latency by using one Python script — MoonL88537 · 2026-07-22