Shengshu launches real-time video model S2: 720p avatars and live video editing at 25-42 FPS, plus spatial video for headsets

赛博禅心 · wechat · 2026-09-15

Shengshu Tech (Vidu) launched real-time video model S2, upping output to 720p at 25-42 FPS, with two models available now via online demo and API.

S2-Avatar (real-time digital human):

S2-Editing (live video editing):

Real-time spatial video: generates slightly different views per eye for headsets — monocular input is generated then converted to stereo; stereo input is stitched horizontally, edited jointly, then split.

Official benchmarks (StreamAV-Bench, Sparkle-Bench, OpenVE, RefVIE, ViViD) show domain SOTA. Project led by Zhang Jintao, Prof. Zhu Jun's PhD student and lead of streaming video generation. Tech report: arxiv.org/pdf/2609.11638

Related event: Shengshu Releases Realtime Video Model Vidu S2(2 posts)→

Original post →

More from Embodied

Embodied channel →