Video DeltaNet hybrid attention speeds livestream video generation 14.5x on 8x B200

Haocheng Xi · hf · 2026-09-18

Video DeltaNet (VDN) combines local softmax attention with a bidirectional linear memory branch for long-range video context, introducing Video Delta Attention (VDA) that updates memory once per frame with spatial tokens. Staged teacher alignment injects the pathway into pretrained models; instantiated on MiniMax H3, video-to-video interactions use the hybrid while text/audio keep softmax. With 8-step distillation and an optimized SGLang serving stack, VDN-H3 denoises a 14.3-second 768p video in 6.70 seconds on eight NVIDIA B200 GPUs—a 14.5x speedup over the 50-step dense H3 baseline.

Original post →

More from Multimodal

Multimodal channel →