Shengshu's Vidu S2 edits live video streams in real time — swap outfits, faces and scenes

量子位 · wechat · 2026-09-15

Shengshu Technology released Vidu S2, just 69 days after S1 and fully available on day one. Three headline upgrades:

Tech: Self-Replay Forcing curbs error drift in streaming generation; a Backbone-Refiner two-stage architecture with async pipelining keeps 720p frame rates; redesigned positional encodings align reference-image injection with timestamps; a VLM agent tracks scene state to plan generation; DPO and Streaming NFT handle RL in the bidirectional and streaming stages. Led by Zhang Jintao, PhD student of Tsinghua's Zhu Jun; technical report available.

Related event: Shengshu launches Vidu S2: real-time interactive and editable video generation(5 posts)→

Original post →

More from Multimodal

Multimodal channel →