ShotStream: Real-Time Multi-Shot Video Generation
jiqizhixin · x · 2026-07-19
Researchers from CUHK and Kuaishou introduced ShotStream, a streaming multi-shot video generation model designed for interactive storytelling.
The core approach uses causal generation to predict the next shot based on historical footage, paired with a dual-memory cache to maintain coherence. It also employs distillation techniques to suppress error accumulation in long sequence generation. The authors claim this method achieves sub-second latency and 16 FPS on a single GPU, reaching a real-time performance balance that long-range autoregressive and bidirectional multi-shot models struggle to achieve simultaneously.
More from Multimodal
- Invideo launches agent-driven video editor that executes edits from plain descriptions — azed_ai · 2026-09-11
- YuE2 music generation gets native ComfyUI support via new PR — LatentSpacer · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11
- Scottish man strolling through his castle: the AI video everyone is sharing — EternalSnow05 · 2026-09-11
- One prompt, full UGC ad: Kling MCP turns a product idea into ready-to-post video — SimplyAnnisa · 2026-09-11
- A Seedance 2.5 quick-start prompt with GPT Image 2.5 hacks — techhalla · 2026-09-11