Ant Group Open-Sources Robotics Video Model

赛博禅心 · wechat · 2026-07-09

Ant's Lingbo team has open-sourced LingBot-Video, a video generation model designed for robotics. It aims to combine internet videos with massive embodied data to generate visual inputs better suited for robot training. Unlike standard video generation that prioritizes visual consistency and aesthetics, robotics training demands physical correctness, including inertia, materials, and action structures.

The post details the model's scale and training pipeline: it features a MoE architecture with 30B total parameters and 3B active parameters for fast generation. The training data includes over 70,000 hours of embodied data from real machines, simulations, and open-source datasets, covering robotic arms, humanoids, and quadrupeds. Training involved multiple stages, progressing from pure images to images and videos, and finally high-resolution refinement. Multi-model reinforcement learning was applied for filtering and correction across perception, physics, and execution dimensions.

Related event: Ant Group Open-Sources LingBot-Video for Embodied AI(26 posts)→

Original post →

More from Embodied

Embodied channel →