Video Foundation Model Trained on Robotic Data
heyshrutimishra · x · 2026-07-13
Unlike models that learn primarily from internet video, LingBot-Video is trained extensively on robotic data.
The author highlights that its training data includes over 70,000 hours of embodied footage, covering robotic manipulation, navigation, and first-person interaction. Compared to merely observing visual appearance, robotic data better teaches the model "how objects move," "how they respond to force," and "what happens during real-world interaction."
Its reward system is also unique: only 1 of 6 signals directly optimizes for visual quality, while the rest enforce constraints on alignment, motion, humanoidal consistency, and physical plausibility.
It achieved a score of 0.620 on RBench, which the post claims is the highest among both open-source and closed-source models on the leaderboard, leading in Manipulation, Long-horizon, and Quadruped tasks.
The author positions it not merely as a video generator, but as a foundation model for embodied AI. It can be used for data generation, physical scene simulation, and exploring robotic policy workflows, ultimately reducing the need for expensive real-world robot trials.
Related event: Embodied Video Model LingBot-Video Goes Open Source(4 posts)→
More from Embodied
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11
- MKBHD goes hands-on with the first folding iPhone; $2,000 48MP selfie cam mocked — alexmacgregor__ · 2026-09-11
- NTU spin-off Ropedia launches HOMIE Gen 2 wearable system to train robots from human experience — liuziwei7 · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- Replit Agent in a robot builds and publishes websites autonomously via MCP — amasad · 2026-09-11