Shengshu Tech on the Path From World Models to Robotics

生数科技 · wechat · 2026-07-19

In a WAIC-related interview, Zhu Jun systematically introduced Shengshu Tech's roadmap from **video generation** to **world models** and **robotic actions**: ViduS1 emphasizes real-time interaction, allowing users to continuously influence character generation via voice. Subsequently, the company linked video prediction, future state modeling, and action generation together, attempting to shift models from "generating content" to "understanding and acting upon the physical world." The article highlighted their **Data Pyramid** and unified **MoT** architecture: first learning world dynamics through massive amounts of general video, first-person video, simulation data, and robot trajectories, then using unified understanding/prediction/action experts for collaborative modeling to reduce error accumulation caused by module fragmentation. The text also mentioned that Motubrain achieved an average success rate of over 95% on more than 50 complex tasks in randomized environments in RoboTwin2.0, which the author views as a phased validation of the general world model approach.

Original post →

More from Embodied

Embodied channel →