LingBot-VLA 2.0 Whole-Body Robot Policy

rohanpaul_ai · x · 2026-07-12

LingBot-VLA 2.0 tackles the challenge of adapting a single policy to various robot embodiments. It scales a single policy across 20 robot configurations, covering arms, grippers, dexterous hands, heads, waists, and mobile bases, using a unified 55-dimensional action format for control.

For training, it first filters out low-quality segments—such as jitter, signal errors, camera mismatches, blurriness, dropped frames, and prolonged stillness—from approximately 90,000 hours of raw robot data, ultimately retaining 50,000 hours of high-quality real robot data. Human videos are also filtered for hand-object interactions, followed by the reconstruction of camera movements and hand poses as action data.

Methodological additions include:

On the Agilex GM-100, it achieves 66.2% progress / 34.4% success, outperforming pi0.5's 59.1% / 32.2%. It also maintains an overall lead on long-horizon mobile tasks, though it is still surpassed by other models on specific individual tasks.

Related event: LingBot-VA/VLA 2.0 Released: Native Embodied Foundation Model(24 posts)→

Original post →

More from Embodied

Embodied channel →