LingBot-Video: Open-Source Video World Model

Savings-Display5123 · reddit · 2026-07-09

LingBot-Video is an open-source video diffusion Transformer featuring a sparse MoE architecture. It has a total parameter count of 13B with 1.4B activated, and incorporates action-conditioned world model capabilities. The model has also undergone multiple rounds of reward-based reinforcement learning training, integrating physical feasibility rewards to support the prediction of robotic rollouts based on actions and hand poses.

Related event: Ant Group Open-Sources LingBot-Video for Embodied AI(26 posts)→

Original post →

More from Multimodal

Multimodal channel →