Video DeltaNet open-sources hybrid-attention VDN-H3, 14.4s video in 11.2s on 8 B200s
BigWideBaker · reddit · 2026-09-05
The team released VDN-Minimax-H3 (VDN-H3), a hybrid-attention video generation model built on MiniMax H3 that generates video faster than it plays.
Key points:
- Fast inference: a 14.4-second clip in 11.23 seconds with 8 denoising steps on 8 B200 GPUs
- Hybrid architecture: a frame-wise linear attention branch for efficiency plus a softmax branch preserving the backbone's visual quality and consistency
- Plug-and-play: adds a separate linear attention branch and two small LoRA adapters that merge into the backbone at inference, leaving backbone weights untouched
- Fully open-source: weights, optimized inference stack, and training code all released
More from Multimodal
- Suno V6 music generation model is about to launch — op7418 · 2026-09-05
- pyannote speaker-diarization-community-1 trends on Hugging Face — pyannote · 2026-09-05
- Creator builds an interactive 90s sitcom genie using fal's H3 Max video model — gorkem · 2026-09-05
- Orbis paper reframes video generation as a steerable, continuous visual process — rohanpaul_ai · 2026-09-05
- Pocket FM killed its business model at $200M ARR to bet everything on AI content — thisguyknowsai · 2026-09-05
- Intangible ships Motion Paths: draw a line in 3D and subjects follow it exactly — azed_ai · 2026-09-05