Open-Source Embodied Video Foundation Model

nikola_mr64990 · x · 2026-07-09

Robbyant has open-sourced LingBot-Video, an embodied intelligence video foundation model utilizing a MoE architecture with 30B total parameters, of which only 3B are activated during inference. The model builds on large-scale internet video pre-training, supplemented with 70,000 hours of embodied data. The post claims it outperforms Wan2.6, Seedance 1.5 Pro, and Cosmos3 Super on RBench, aiming to accelerate the deployment of embodied applications like robotics.

Related event: Ant Group Open-Sources LingBot-Video for Embodied AI(26 posts)→

Original post →

More from Embodied

Embodied channel →