Shengshu argues world models will bridge digital content and physical action
生数科技 · wechat · 2026-07-21
At WAIC 2026, Shengshu Technology and Wondershare hosted a forum on “world-model-driven AI film production.” In his talk, Tsinghua professor and Shengshu founder Zhu Jun argued that AI is moving from generating digital content to understanding, predicting, and acting in the physical world.
Key points
- World models as infrastructure: He framed world models as the bridge between digital and physical worlds, with three core capabilities: understanding, imagination/prediction, and action.
- Data pyramid: Shengshu’s approach relies on a multi-layer data stack, from internet video to first-person human action video and robot operation data, with video as the main carrier.
- Unified architecture: He described a “MoT” unified architecture that combines understanding, generation, and action experts in one model.
- Video to real-time interaction: The company’s video models can do real-time, streaming generation, support voice-driven interaction, and output up to 540p at 42 FPS.
- Embodied intelligence: A newer model can drive heterogeneous robots, decompose text instructions into actions, and use future-state prediction to improve planning and stability.
The overarching thesis is that world models are not just an upgrade to video or language models, but a new paradigm for content production and robotics.
Related event: Shengshu Technology Highlights World Model as New AI Paradigm(2 posts)→
More from Multimodal
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11
- Skyfall GS Uses Flux to Refine Gaussian Splatting, Accepted at ECCV 2026 — ducha_aiki · 2026-09-11