vLLM-Omni makes MiniMax H3 real-time: 10s MP4 in 8.7s on 8x B300, 49 DiT passes cut to 4
机器之心 · wechat · 2026-09-13
The vLLM-Omni team details full-pipeline serving optimization for MiniMax H3's joint audio-video generation, covering the Qwen3-VL encoder, joint DiT, VAEs, cross-process transfer, and H.264/AAC muxing.
On 8x NVIDIA B300, FastH3 cuts the denoising loop from 49 DiT forward passes to 4, producing a 10.125s MP4 in 8.678–8.710s — real-time by full-response readiness. Lossless system-level optimizations (TRTLLMATTN, FastUlysses symmetric-memory comm, fused DiT ops, 8-GPU parallel VAE decode, GPU-side uint8 output prep) yield a 30.8% latency reduction (1.445x) vs Diffusers. Optional scaling paths include Distributed Layerwise Offload (37.5% HBM savings at 5.1% latency cost), online FP8 (38.9% lower peak HBM), and SAGE+Skip-Softmax attention (1.237x), with evidence lines rigorously separated.
More from Infra
- Fiber-to-Chip Coupling Needs Path Analysis, Not Single Metrics Like 1dB Loss — jwt0625 · 2026-09-13
- 100% GPU Utilization Can Still Be Slow: A 42-Page Handbook on SMs, Warps and Memory — techNmak · 2026-09-13
- Pentagon reportedly weighs $5B loan to Fluidstack for AI infrastructure supply chain — VraserX · 2026-09-13
- AMD overtakes Qualcomm to become world's third-largest fabless chip company — xiaosun86 · 2026-09-13
- vLLM Dev Pushes Back as 5 Specialized Inference Engines Launch in One Month — AccBalanced · 2026-09-13
- Ex-compute veteran rebuts Dario's 'pace the frontier': compute doesn't idle, it reroutes — basedjensen · 2026-09-13