Shengshu Tech Unveils 5-Stage Roadmap for General World Models
生数科技 · wechat · 2026-08-24
Zhu Jun, founder of Shengshu Tech, outlined a 5-stage evolutionary roadmap (L1-L5) for General World Models at the World Robot Conference, emphasizing a closed-loop system of understanding, prediction, and action.
Core Definition
A general world model integrates understanding, prediction, and action into a coupled feedback system, rather than acting as a standalone generator or policy model.
Technical Pillars
- Data: A multi-layer pyramid ranging from web videos to robot interaction data, utilizing "imperfect data" (failures/corrections) for learning.
- Architecture: MoT (Mixture-of-Transformers) unifies image, video, language, and robotic action processing.
- Compute: Balances large-scale pre-training with real-time inference acceleration.
5-Stage Roadmap
- L1 World Generation: Achieved via the Vidu video model.
- L2 Interactive World: ViduS1 enables real-time voice interaction.
- L3 Actionable World: Motubrain unifies environmental understanding with action generation, topping the RoboTwin2.0 benchmark and verified on nearly 10 robot types.
- L4 Autonomous World Agent: Capable of long-term planning and active exploration.
- L5 World Orchestrator: Coordinates multi-agent collaboration for complex tasks.
Future work for L4/L5 includes overcoming six challenges, such as learning physical laws, persistent memory, and online learning.
More from Embodied
- Matic robot vacuum mop review: Cleans entire apartment in an hour — yungcontent · 2026-08-24
- Slow-Mo Look at Humanoid Robot's 400m Champion Running Form — herbiebradley · 2026-08-24
- Local DeepSeek V4 Flash Benchmark: 24 tok/s on Epyc + RTX 5090 — IntravenusDeMilo · 2026-08-24
- R1 device now runs local agents like Hermes/Claude Code; OS3 coming soon — jesselyu · 2026-08-24
- SWANC consumer robots open pre-orders; 80 humanoids spell "BEIJING" autonomously — 创业邦 · 2026-08-24
- A humanoid joint repair costs $4K+, robot repair becomes a livestream niche — 创业邦 · 2026-08-24