LingBot-VLA 2.0 Whole-Body Robot Policy
rohanpaul_ai · x · 2026-07-12
LingBot-VLA 2.0 tackles the challenge of adapting a single policy to various robot embodiments. It scales a single policy across 20 robot configurations, covering arms, grippers, dexterous hands, heads, waists, and mobile bases, using a unified 55-dimensional action format for control.
For training, it first filters out low-quality segments—such as jitter, signal errors, camera mismatches, blurriness, dropped frames, and prolonged stillness—from approximately 90,000 hours of raw robot data, ultimately retaining 50,000 hours of high-quality real robot data. Human videos are also filtered for hand-object interactions, followed by the reconstruction of camera movements and hand poses as action data.
Methodological additions include:
- Sparse MoE: Allows different action tokens to use a small number of specialized networks while sharing experts to retain general skills
- Depth/video prediction auxiliary objectives: Forces the model to understand object geometry and scene changes during the next action segment
On the Agilex GM-100, it achieves 66.2% progress / 34.4% success, outperforming pi0.5's 59.1% / 32.2%. It also maintains an overall lead on long-horizon mobile tasks, though it is still surpassed by other models on specific individual tasks.
Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→
More from Embodied
- The Humanoid AI raises $152M Series A at a $1.35B valuation — RazRazcle · 2026-07-22
- SceniX joins World Labs to close the real-to-sim gap for robot learning — davidyin44 · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- A quadruped robot gets a custom glow-up with a new shell and screen — DynamicWebPaige · 2026-07-22
- NVIDIA says physical AI starts in simulation with OpenUSD and synthetic data — MonaJalal_ · 2026-07-22