vLLM ships day-0 support for Qwen-Image-2.1 with cross-step prefix KV cache

Alibaba_Qwen · x · 2026-09-20

vLLM announced day-0 support for Qwen-Image-2.1, which Alibaba's Qwen team thanked them for. The model pairs a 7.1B single-stream DiT (block-causal attention) with a Qwen3-VL-8B text encoder and a 16x RGBA autoencoder, serving both text-to-image and editing in one pipeline. vLLM-Omni accelerates it via cross-step prefix KV cache reuse, dedicated CUDA Graphs, continuous batching, tensor/Ulysses sequence parallelism, FP8 weights and CPU offloading. Recommended defaults: 40 steps, cfg-scale 1.0.

Original post →

More from Infra

Infra channel →