RunningHub open-sources H3Lightning, speeding up MiniMax H3 video generation 12x
智东西 · wechat · 2026-09-11
- RunningHub open-sourced H3Lightning, cutting MiniMax H3's 5-second video generation from 348.8s to 28.7s — roughly 12x faster, 92% less waiting.
- Optimization stack: distilled model cutting steps from 50 to 9 (5.8x alone), plus SageAttention2, Cache-DiT, torch.compile, and a TP2+Ulysses4 parallel scheme for NVLink-free PCIe clusters (12% faster, 14GiB less VRAM).
- Integrated into SGLang's multimodalgen engine, tested on 8x RTX6000D, keeping BF16 precision; full deployment guide on GitHub.
- Over 1,000 creators have published 10,000 H3 workflows on RunningHub; its enterprise API handles tens of millions of daily calls.
More from Infra
- Qwen3-8B gets a KV-approximation add-on that halves prefill time without touching the model — teortaxesTex · 2026-09-11
- Burning through two ChatGPT resets a day, user coins the "Huang-Altman Law" — yihui_indie · 2026-09-11
- OpenRouter agents now out-consume humans as AI usage arrives in three waves — AccBalanced · 2026-09-11
- Nvidia Is Now Core to Every Major Robotaxi Stack at Commercial Scale — pdamodaran · 2026-09-11
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11