vLLM-Omni: serving voice, video, and diffusion models explained by Red Hat AI
vllm_project · x · 2026-09-15
At vLLM Office Hours #57, Alex Brooks, Ricardo Noriega, and Nick Cao of Red Hat AI walked through the vLLM-Omni architecture and recent updates. Unlike text models that advance one token at a time, voice and video-with-audio models need entirely different scheduling. A new recap article, video, and slides are available.
More from Infra
- Expert-lookahead delivers 10%+ speedup for MoE inference on low-memory Macs — carloslfu · 2026-09-15
- Local voice assistant keeps its memory across model swaps, thanks to open-source Mnemosyne — Enough_Leopard3524 · 2026-09-15
- Hugging Face reportedly acquired by NVIDIA — Turn_Trout · 2026-09-15
- NextDC accused of using AI to draft lobbying letters, with AI errors, for 225MW datacenter expansion — nordicinst · 2026-09-15
- Running Z-Image Turbo locally on an RX 6800: full ROCm setup, 43s per image — AstroFieldsGlowing · 2026-09-15
- Sparse GEMM Deserves Attention: Insights From a Dense-GEMM Optimizer — goyal__pramod · 2026-09-15