Video Foundation Model Trained on Robotic Data

heyshrutimishra · x · 2026-07-13

Unlike models that learn primarily from internet video, LingBot-Video is trained extensively on robotic data.

The author highlights that its training data includes over 70,000 hours of embodied footage, covering robotic manipulation, navigation, and first-person interaction. Compared to merely observing visual appearance, robotic data better teaches the model "how objects move," "how they respond to force," and "what happens during real-world interaction."

Its reward system is also unique: only 1 of 6 signals directly optimizes for visual quality, while the rest enforce constraints on alignment, motion, humanoidal consistency, and physical plausibility.

It achieved a score of 0.620 on RBench, which the post claims is the highest among both open-source and closed-source models on the leaderboard, leading in Manipulation, Long-horizon, and Quadruped tasks.

The author positions it not merely as a video generator, but as a foundation model for embodied AI. It can be used for data generation, physical scene simulation, and exploring robotic policy workflows, ultimately reducing the need for expensive real-world robot trials.

Related event: Embodied Video Model LingBot-Video Goes Open Source(4 posts)→

Original post →

More from Embodied

Embodied channel →