LingBot-Video Training: 6D Physical Rewards and Real Robot Data

thetripathi58 · x · 2026-07-09

LingBot-Video changes the traditional scoring mechanism of video models that only pursue visual aesthetics. It adopts a single-step GRPO algorithm and introduces 6 precise reward signals: visual quality, image-text alignment, dynamics, motion coherence, human action consistency, and physical plausibility, ensuring the model understands real physical laws.

To prevent the model from "faking" physics, the team trained it using over 70,000 hours of real robot data, including real-hand operations, walking trajectories, and first-person perspectives from humanoid and quadruped robots.

Related event: Ant Group Open-Sources LingBot-Video for Embodied AI(26 posts)→

Original post →

More from Embodied

Embodied channel →