FlashForward reuses in-flight KV cache to speed long video diffusion up to 2.9x
Yikai Wang · hf · 2026-09-29
FlashForward eliminates cache-update-only forwards in few-step autoregressive video diffusion by directly reusing the in-flight KV computed during each chunk's denoising, and assigns one GPU per stage so chunks occupy different stages concurrently.
Key points:
- The noisy stage-matched history causes appearance/motion drift, so sparse clean anchor latents are generated beforehand as two-sided conditioning — anchors give coarse long-range structure, dense stage-matched history preserves recent detail;
- With up to 4 GPUs it runs 1.16–1.69x faster than HiAR and 1.42–2.92x faster than Self-Forcing for 20s+ videos at 480p/720p across 1.3B and 14B backbones;
- Higher VBench scores at 480p with stability across 20s, 35s and 65s durations.
More from Multimodal
- Full prompt released: 30s photoreal selfie vlog with Seedance 2.5 — techhalla · 2026-09-29
- Seedance 2.5 single-prompt video turns the Shire into a gritty selfie vlog — techhalla · 2026-09-29
- Match cuts with generative AI: transform elements mid-action while keeping motion continuity — ArcaArtificial · 2026-09-29
- invideo launches agentic video editor that edits a real multitrack timeline — azed_ai · 2026-09-29
- Redditor Iterates a Full Show Episode with Opus 5.5 — No One-Shot, but Pretty Happy — weakcper · 2026-09-29
- MiniMax H3 used to create animated Lord of the Rings demo — azed_ai · 2026-09-29