Shengshu lays out a world-model stack for video generation and embodied AI
生数科技 · wechat · 2026-07-21
At WAIC 2026, Shengshu Technology and Wondershare held a forum on “a new paradigm for AI film production driven by world models.” In a long keynote, founder Zhu Jun said AI is evolving from content generation to a system that can understand the world, predict what will happen next, and act in it.
The argument
- World models as a new infrastructure: Zhu framed world models as the next step after language and multimodal models, with the goal of linking digital and physical worlds.
- Three required abilities: A machine world model should understand the environment, predict the effects of future actions, and learn while acting.
- Video-first data strategy: Shengshu’s training stack is built around a six-layer data pyramid, with internet video, first-person human action video, and robot operation data as key ingredients.
- One model, multiple outputs: The company’s MoT unified architecture combines understanding, generation, and action experts in one system, allowing the same base model to output video or robot actions depending on the task.
- Real-time video generation: Shengshu says its streaming model supports voice instructions, character definition, unlimited-length dialogue, and outputs up to 540p at 42 FPS.
- Embodied AI: The latest model can control multiple robot embodiments, break down text tasks such as arranging flowers or watering plants, and predict future states to improve planning and safety.
- Efficiency gains: Zhu said the newer video foundation model improves cross-scenario and cross-task generalization, learning better skills with less data than earlier versions.
The talk positions world models as more than an upgraded video model or a language-model extension: Shengshu wants them to become the foundation for both creative production and embodied intelligence.
Related event: Shengshu Technology Highlights World Model as New AI Paradigm(2 posts)→
More from Multimodal
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11