Video Foundation Model Trained on Robotic Data
heyshrutimishra · x · 2026-07-13
Unlike models that learn primarily from internet video, LingBot-Video is trained extensively on robotic data.
The author highlights that its training data includes over 70,000 hours of embodied footage, covering robotic manipulation, navigation, and first-person interaction. Compared to merely observing visual appearance, robotic data better teaches the model "how objects move," "how they respond to force," and "what happens during real-world interaction."
Its reward system is also unique: only 1 of 6 signals directly optimizes for visual quality, while the rest enforce constraints on alignment, motion, humanoidal consistency, and physical plausibility.
It achieved a score of 0.620 on RBench, which the post claims is the highest among both open-source and closed-source models on the leaderboard, leading in Manipulation, Long-horizon, and Quadruped tasks.
The author positions it not merely as a video generator, but as a foundation model for embodied AI. It can be used for data generation, physical scene simulation, and exploring robotic policy workflows, ultimately reducing the need for expensive real-world robot trials.
Related event: Embodied Video Model LingBot-Video Goes Open Source(4 posts)→
More from Embodied
- Tesla Robotaxi expands to Orlando and Tampa, reaching seven U.S. markets — XFreeze · 2026-07-22
- Hands-on robotics workshop on Saturday may be the last in-person session before August — StewartalsopIII · 2026-07-22
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21
- Gritt says an 8-person crew now installs 3,000 to 4,000 solar panels a day — HaktanSuren · 2026-07-21
- A helium-powered flying robot whale aims to be a quiet companion pet — chris_j_paxton · 2026-07-21