RunningHub open-sources a 12x speedup for MiniMax H3 video generation, keeping full BF16 precision
量子位 · wechat · 2026-09-10
RunningHub (backed by Haila Cloud) open-sourced MiniMax-H3-MultiGPU-Lightning, a full inference-acceleration stack for the open-source H3 video model.
- Results: Generating a 5s 1344×768 video drops from 348.8s (BF16, 50 steps, 4x RTX 6000D) to 28.7s — roughly 12x faster with 92% less time, while preserving full BF16 precision. 15s dual-reference videos run in 48–73s on 8 cards.
- The three cuts: (1) a post-trained RH acceleration model reduces steps from 50 to 4–9 (8 for fast motion), an 8x win alone; (2) SageAttention2, Cache-DiT caching, and torch.compile further squeeze the remaining compute; (3) on NVLink-less PCIe setups, benchmarking picked TP2+Ulysses4 parallelism (12% faster, 14GiB less VRAM than TP4+Ulysses2), all orchestrated via SGLang multimodalgen.
- Local-deployable: The stack targets commodity pro GPUs rather than datacenter hardware; install, download, service launch, and inference test docs are public for studios and small teams.
- RunningHub previously released H3 ComfyUI nodes; creators have built 10,000 workflows for e-commerce video, digital humans, and comics.
More from Multimodal
- Creator makes 2D electro-pop anime music video with just a prompt using MiniMax H3 — Hailuo_AI · 2026-09-11
- Street View to driving footage: GPT Astra fetches images, MiniMax H3 turns them into dashcam video — Hailuo_AI · 2026-09-11
- Single-author ECCV 2026 paper makes rolling shutter correction practical — ducha_aiki · 2026-09-11
- AI digital human covers Japanese classic so realistically viewers can't tell — JourneymanChina · 2026-09-11
- ComfyUI Style Explorer Adds LoRA Preview Catalog and Sharing — neonsparksuk · 2026-09-11
- 4 favorite Midjourney V6.1 --sref style codes, ready to copy — michaelrabone · 2026-09-11