vLLM ships day-0 support for Qwen-Image-2.1 with cross-step prefix KV cache
Alibaba_Qwen · x · 2026-09-20
vLLM announced day-0 support for Qwen-Image-2.1, which Alibaba's Qwen team thanked them for. The model pairs a 7.1B single-stream DiT (block-causal attention) with a Qwen3-VL-8B text encoder and a 16x RGBA autoencoder, serving both text-to-image and editing in one pipeline. vLLM-Omni accelerates it via cross-step prefix KV cache reuse, dedicated CUDA Graphs, continuous batching, tensor/Ulysses sequence parallelism, FP8 weights and CPU offloading. Recommended defaults: 40 steps, cfg-scale 1.0.
More from Infra
- SpaceX's orbital AI data centers weigh up to 4,000 kg each, filing seeks 1M satellites — XFreeze · 2026-09-20
- Local Models for Personal Agents: GPT Luna Surprises a Coding-Agent Veteran — gized00 · 2026-09-20
- OpenAI hardware VP details first custom chip Jalapeño and its nine-month tape-out — bigdata · 2026-09-20
- focus-llama: a llama.cpp fork implementing Declarative Attention for up to 0.71x decode time — Ok-Shower7286 · 2026-09-20
- VTrain, a Vulkan-based resident trainer, fixes memory leak and offloads more work to GPU — Savantskie1 · 2026-09-20
- Dev builds helmstudio, an MLX-first local launcher for open models on Apple Silicon — janishar · 2026-09-20