Ant Group Open-Sources Robotics Video Model
赛博禅心 · wechat · 2026-07-09
Ant's Lingbo team has open-sourced LingBot-Video, a video generation model designed for robotics. It aims to combine internet videos with massive embodied data to generate visual inputs better suited for robot training. Unlike standard video generation that prioritizes visual consistency and aesthetics, robotics training demands physical correctness, including inertia, materials, and action structures.
The post details the model's scale and training pipeline: it features a MoE architecture with 30B total parameters and 3B active parameters for fast generation. The training data includes over 70,000 hours of embodied data from real machines, simulations, and open-source datasets, covering robotic arms, humanoids, and quadrupeds. Training involved multiple stages, progressing from pure images to images and videos, and finally high-resolution refinement. Multi-model reinforcement learning was applied for filtering and correction across perception, physics, and execution dimensions.
Related event: Ant Group Open-Sources LingBot-Video for Embodied AI(26 posts)→
More from Embodied
- Nothing phone mockup turns a film joke into a modular design meme — ZeYanjie · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- Humanoid robot sorting packages in a warehouse sparks debate over job loss — MonaJalal_ · 2026-07-22
- NVIDIA pushes OpenUSD as the common layer for simulation and physical AI — MonaJalal_ · 2026-07-22
- A quadruped robot gets a custom glow-up with a new shell and screen — DynamicWebPaige · 2026-07-22
- A VR teleop demo for an SO-101 arm gets absurdly low latency by using one Python script — MoonL88537 · 2026-07-22