vLLM-Omni Enables Real-Time Serving for MiniMax H3
vllm_project · x · 2026-09-02
The vLLM team blog post details system-wide optimizations enabling real-time video generation for MiniMax H3. Key improvements include Fast Ulysses communications for long-context attention, fully fused DiT operators (RMSNorm+RoPE, SwiGLU), parallel video/audio VAE decoding across 8 GPUs, parallel H.264/AAC muxing, and integration with FastVideo's 4-step FastH3 distillation. On an 8xB300 GPU setup, generating a 10.125-second MP4 takes just 8.7 seconds, achieving real-time performance.
More from Infra
- Morgan Stanley Raises Google TPU Sales Forecast to $108B by 2028 — Beth_Kindig · 2026-09-02
- Anyscale Acquired by AI Cloud Nscale for $1.65B — jfiance · 2026-09-02
- DIY CUDA box with unlocked CMP170HX mining cards hits 4000 tps prompt processing for Qwen Flash Next — Miserable-Dare5090 · 2026-09-02
- Broadcom's VMware Private AI Push Signals Enterprise AI Moving From Pilots to Production — DavidLinthicum · 2026-09-02
- US holds majority of global compute, maintaining massive advantage — peterwildeford · 2026-09-02
- ARK analyst: every dollar of GDP per capita needs 1 kWh per capita — skorusARK · 2026-09-02