SVEET: streaming video editing with a diffusion model hits 15 FPS on a single H100
SJTU · hf · 2026-09-22
Researchers from SJTU propose SVEET, a streaming video editing framework built on a pretrained bidirectional video diffusion model, enabling high-quality auto-regressive editing in real time.
Key ideas
- Systematically revisits existing video-to-video diffusion approaches and distills two principles for streaming adaptation: backbone feature disentanglement and conditional frame independence
- An auxiliary branch encodes the source video with temporally independent self-attention, injecting intermediate features into corresponding backbone blocks for streaming-compatible control
- A decoupled training scheme enforces orthogonality between video controllability and model causality, keeping the two objectives compatible at inference and enabling zero-shot transfer across heterogeneous backbones
Results
- Superior editing quality with 15 FPS on a single H100, no auxiliary acceleration needed
- Code: https://github.com/YujiaHu1109/SVEET
More from Multimodal
- MiniMax H3 Turbo LoRAs plagued by high-pitch audio screeches in loud scenes — hurdurdur7 · 2026-09-22
- AI-generated NFS: Hot Pursuit-style chase video wows X, full prompt shared — CurieuxExplorer · 2026-09-22
- Redditor uses Wan and Hailuo to make an AI fanedit of Conan the Destroyer with new scenes — Itchy-Advertising857 · 2026-09-22
- Reddit User Turns Nano Banana 2 Into a Free Character Creator via Chart-Based Prompting — Extension-Fee-8480 · 2026-09-22
- Seedance 2.5 adds Draft mode: 480P previews then 1080P finals, up to 77% cheaper — 火山引擎 · 2026-09-22
- Where ComfyUI v0.37 Hides Its Startup Arguments (and Why You No Longer Need to Set Memory Manually) — tostane · 2026-09-22