vLLM-Omni makes MiniMax H3 real-time: 10s MP4 in 8.7s on 8x B300, 49 DiT passes cut to 4

机器之心 · wechat · 2026-09-13

The vLLM-Omni team details full-pipeline serving optimization for MiniMax H3's joint audio-video generation, covering the Qwen3-VL encoder, joint DiT, VAEs, cross-process transfer, and H.264/AAC muxing.

On 8x NVIDIA B300, FastH3 cuts the denoising loop from 49 DiT forward passes to 4, producing a 10.125s MP4 in 8.678–8.710s — real-time by full-response readiness. Lossless system-level optimizations (TRTLLMATTN, FastUlysses symmetric-memory comm, fused DiT ops, 8-GPU parallel VAE decode, GPU-side uint8 output prep) yield a 30.8% latency reduction (1.445x) vs Diffusers. Optional scaling paths include Distributed Layerwise Offload (37.5% HBM savings at 5.1% latency cost), online FP8 (38.9% lower peak HBM), and SAGE+Skip-Softmax attention (1.237x), with evidence lines rigorously separated.

Original post →

More from Infra

Infra channel →