vLLM-Omni v0.22 adds Cosmos 3, robot serving, and broader quantization support
MonaJalal_ · x · 2026-07-22
vLLM-Omni v0.22.0 is a major upgrade focused on production-grade multimodal serving.
It adds day-0 support for NVIDIA Cosmos 3 world models, robot serving via DreamZero + OpenPI realtime API, production TTS for models like Qwen3-TTS and Qwen3-Omni, faster image/video/diffusion support including Wan 2.2, HunyuanVideo 1.5, and LTX-2.3, plus broader quantization and hardware coverage such as FP8/INT8, MXFP4/MXFP8, and W4A16. The release notes mention 339 commits, 124 contributors, and 52 new contributors.
More from Infra
- BloombergNEF: U.S. data centers could use 20% of national electricity by 2035 — coinfanking · 2026-07-22
- Reddit user runs Nemotron Ultra 550B across aging MI50 and P40 GPU rigs — Old_Grapefruit8774 · 2026-07-22
- SK Hynix denies talks to buy Intel’s Ohio fab after market rumors — rwang07 · 2026-07-22
- Ineffable Labs takes delivery of a Vera Rubin NVL72 cluster — deanwball · 2026-07-22
- Chemical Giant Buys Its Way Into the GPU Rack Market — shashib · 2026-07-22
- GPT-6 is said to be near, with OpenAI betting on faster inference and custom chips — haider1 · 2026-07-22