Qwen-Image 2.1 runs on SGLang: image generation in 18.7s on a single RTX 4090

Alibaba_Qwen · x · 2026-09-21

SGLang-Diffusion announced day-0 support for Alibaba's Qwen-Image 2.1, acknowledged by the Qwen team.\n\nPerformance (no quantization, 40 denoising steps, warmed HTTP latency including PNG output):\n- RTX 4090 24GB with CPU offload: 1024×1024 generation in 18.7s, editing in 21.7s, 22.7 GiB peak GPU memory\n- RTX PRO 6000 96GB: 8.0s generation, 9.6s editing\n\nA single checkpoint covers text-to-image, multi-image editing, and transparent RGBA output, with native TP/SP inference, LoRA, and OpenAI-compatible APIs.

Original post →

More from Infra

Infra channel →