Shengshu argues world models will bridge digital content and physical action
生数科技 · wechat · 2026-07-21
At WAIC 2026, Shengshu Technology and Wondershare hosted a forum on “world-model-driven AI film production.” In his talk, Tsinghua professor and Shengshu founder Zhu Jun argued that AI is moving from generating digital content to understanding, predicting, and acting in the physical world.
Key points
- World models as infrastructure: He framed world models as the bridge between digital and physical worlds, with three core capabilities: understanding, imagination/prediction, and action.
- Data pyramid: Shengshu’s approach relies on a multi-layer data stack, from internet video to first-person human action video and robot operation data, with video as the main carrier.
- Unified architecture: He described a “MoT” unified architecture that combines understanding, generation, and action experts in one model.
- Video to real-time interaction: The company’s video models can do real-time, streaming generation, support voice-driven interaction, and output up to 540p at 42 FPS.
- Embodied intelligence: A newer model can drive heterogeneous robots, decompose text instructions into actions, and use future-state prediction to improve planning and stability.
The overarching thesis is that world models are not just an upgrade to video or language models, but a new paradigm for content production and robotics.
Related event: Shengshu Technology Highlights World Model as New AI Paradigm(2 posts)→
More from Multimodal
- Midjourney V8.2 adds personalization and shows off stylized image outputs — Mr_AllenT · 2026-07-27
- Midjourney’s image variety draws a Krea 2 comparison and asks how to reproduce it — diffusion_throwaway · 2026-07-27
- AI short film sets a 1985 dystopia to music and leans into cinema — ProfessorKey98 · 2026-07-27
- A new BOTPD episode made with Google Omni turns into an AI chase-scene parody — ScriptLurker · 2026-07-27
- A new LoRA recreates GTA: San Andreas’ classic RenderWare-era visuals — Humble-Pick7172 · 2026-07-27
- Enabling dynamic VRAM cuts LTX 2.3 video generation to 168s on an AMD R9700 — xdcfret1 · 2026-07-27