Video DeltaNet hybrid attention speeds livestream video generation 14.5x on 8x B200
Haocheng Xi · hf · 2026-09-18
Video DeltaNet (VDN) combines local softmax attention with a bidirectional linear memory branch for long-range video context, introducing Video Delta Attention (VDA) that updates memory once per frame with spatial tokens. Staged teacher alignment injects the pathway into pretrained models; instantiated on MiniMax H3, video-to-video interactions use the hybrid while text/audio keep softmax. With 8-step distillation and an optimized SGLang serving stack, VDN-H3 denoises a 14.3-second 768p video in 6.70 seconds on eight NVIDIA B200 GPUs—a 14.5x speedup over the 50-step dense H3 baseline.
More from Multimodal
- Suno's Pricing Change Frustrates Users, Who Point to Open Alternatives — Bedrovelsen · 2026-09-18
- Pika simplifies AI creation: pick what you want to make, not which model to use — minchoi · 2026-09-18
- Another Qwen image model appears imminent, hints Andrew Carr — andrew_n_carr · 2026-09-18
- UFO: unified omni-condition alignment evaluation for multimodal image generation, +15.25% human correlation — ustc-community · 2026-09-18
- AI-Generated Ming Dynasty Dynasty Reenactment Wows Viewers Online — angadc · 2026-09-18
- Creator builds StyleGAN dataset from ~3,000 synthetic JAX diffusion images — makeitrad1 · 2026-09-18