Shengshu unveils Vidu, ViduS1 and Motubrain as a full world-model stack at WAIC 2026
生数科技 · wechat · 2026-07-24
At WAIC 2026, Shengshu Technology showcased a world-model stack spanning digital generation, real-time interaction, and physical action.
What it presented
- Vidu: a world-generation video model for text-to-video, image-to-video, and reference-to-video workflows, aimed at consistent, physically plausible content creation.
- ViduS1: a real-time interactive video model that can create a personalized character from an initial image and voice, then sustain live conversation with generated expressions, gestures, gaze, and body motion. The post says it supports video-call style output at 540p/25fps, with up to 42fps.
- Motubrain: a world-action model positioned as a general brain for robots, using a Mixture-of-Transformer architecture to unify environment understanding, world-state prediction, and action generation.
Product and ecosystem claims
- Vidu is described as having SaaS, MaaS API, Agent, and ViduClaw product forms, serving advertising, short drama, animation, film, and e-commerce use cases.
- ViduS1 is pitched for AI companionship, virtual idols, interactive livestreaming, NPCs, digital humans, education, customer service, and XR.
- Motubrain is framed as a multi-robot, multi-task, long-horizon control model, and the post claims it ranked first on RoboTwin 2.0 Clean and Randomized with scores of 95.8 and 96.1, and led WorldArena with an EWMScore of 63.77.
Industry angle
The article argues that world models are becoming a new infrastructure layer for AI: not just generating content, but understanding environments, predicting change, and acting in the physical world. It also highlights an outbound content alliance built around AI film/TV production and global distribution.
More from Embodied
- Masked Visual Actions turns 15 hours of robot video into a zero-shot world model — jbhuang0604 · 2026-07-24
- A retro wrist computer becomes a joke about humanoid robot emergency overrides — seanmcdonaldxyz · 2026-07-24
- Vision-language-action systems can still miss the next action chunk — StillThese3747 · 2026-07-24
- OpenAI could become the intelligence layer for dozens of robot brands — VraserX · 2026-07-24
- RealSense launches D585 Pro depth camera for humanoids and AMRs — lukas_m_ziegler · 2026-07-24
- OpenAI’s Codex keyboard ships, and early users say it’s pricey but fun — APPSO · 2026-07-24