Shengshu's Vidu S2 edits live video streams in real time — swap outfits, faces and scenes
量子位 · wechat · 2026-09-15
Shengshu Technology released Vidu S2, just 69 days after S1 and fully available on day one. Three headline upgrades:
- S2-Editing: real-time outfit, character, background and style swaps on a playing video stream — live 'Photoshop for video.' Testing shows solid identity consistency and cloth deformation, with occasional smearing and hat-hair artifacts.
- S2-Avatar: real-time interactive avatars bumped from 540p to 720p, now executing large motions (dancing, turning) on voice command and accepting reference images mid-conversation to hold objects or change clothes.
- Real-time spatial video: streams can be converted to stereo views for VR headsets, treating the physical world as a canvas.
Tech: Self-Replay Forcing curbs error drift in streaming generation; a Backbone-Refiner two-stage architecture with async pipelining keeps 720p frame rates; redesigned positional encodings align reference-image injection with timestamps; a VLM agent tracks scene state to plan generation; DPO and Streaming NFT handle RL in the bidirectional and streaming stages. Led by Zhang Jintao, PhD student of Tsinghua's Zhu Jun; technical report available.
More from Multimodal
- Flam's 26B MoE Falcon model returns first token in 30ms, specialized for Indic languages — testingcatalog · 2026-09-15
- 4DAnyone Falters on Coats and Odd Poses but Remains a Great Layout Tool — mickmumpitz · 2026-09-15
- Creator builds a template library for Gemini + HyperFrames to clone reels on demand — toolstelegraph · 2026-09-15
- Running Z-Image Turbo locally on an RX 6800: full ROCm setup, 43s per image — AstroFieldsGlowing · 2026-09-15
- Seedance 2.0 still beats 2.5 for FPV shots, testers say — better speed and motion — azed_ai · 2026-09-15
- Redditor shares short video made entirely with AI — anotheraccountaus · 2026-09-15