Shengshu launches real-time video model S2: 720p avatars and live video editing at 25-42 FPS, plus spatial video for headsets
赛博禅心 · wechat · 2026-09-15
Shengshu Tech (Vidu) launched real-time video model S2, upping output to 720p at 25-42 FPS, with two models available now via online demo and API.
S2-Avatar (real-time digital human):
- Generates an interactive avatar from one image — talk to it, change clothes, hold objects, swap scenes;
- New "dynamic reference" lets you inject reference images mid-conversation to update role, props, outfits;
- A VLMAgent layer maintains state consistency: when generating follow-up prompts it explicitly preserves prior state ("still holding the cup") while only changing what's asked — video model generates, agent maintains state;
- Training added solo dance and 2D/3D animation data for fuller body-action following; low-res backbone for motion/temporal plus single-step Refiner for appearance;
- Self-Replay Forcing (SRF): re-noises segments of the model's own long generations for causal replay training, reducing error accumulation in continuous generation.
S2-Editing (live video editing):
- Takes a continuous video stream and applies text-driven style transfer, virtual try-on, character swap, or background replacement in real time, with motion supplied by the source video;
- Tackles ghosting with Frame-Aligned Attention: target frames only exchange information with the source frame at the same timestamp, while reference images keep conditioning appearance across frames; reuses SRF plus streaming memory optimization.
Real-time spatial video: generates slightly different views per eye for headsets — monocular input is generated then converted to stereo; stereo input is stitched horizontally, edited jointly, then split.
Official benchmarks (StreamAV-Bench, Sparkle-Bench, OpenVE, RefVIE, ViViD) show domain SOTA. Project led by Zhang Jintao, Prof. Zhu Jun's PhD student and lead of streaming video generation. Tech report: arxiv.org/pdf/2609.11638
Related event: Shengshu Releases Realtime Video Model Vidu S2(2 posts)→
More from Embodied
- AI passport hack: letting an AI model write the entire firmware for a dynamic card — oran_ge · 2026-09-15
- Steam Frame review knocks support cycle; author points to Apple's iPhone 11 getting iOS 27 — AndrewSchmidtFC · 2026-09-15
- Agent as Policy: a general agent writes code to control real robots directly — Mengzhao Jia · 2026-09-15
- Engineer uses AI image gen to create PCB silkscreen art, calls it a life hack — realwillreil · 2026-09-15
- Bemo Technology to list in Hong Kong as direct-drive modules surge 45x in two years — 创业邦 · 2026-09-15
- MinimalPhone2, a QWERTY 'focus phone' with a kill switch, crowdfunds $2.3M — 创业邦 · 2026-09-15