Qwen-Image 2.1 runs on SGLang: image generation in 18.7s on a single RTX 4090
Alibaba_Qwen · x · 2026-09-21
SGLang-Diffusion announced day-0 support for Alibaba's Qwen-Image 2.1, acknowledged by the Qwen team.\n\nPerformance (no quantization, 40 denoising steps, warmed HTTP latency including PNG output):\n- RTX 4090 24GB with CPU offload: 1024×1024 generation in 18.7s, editing in 21.7s, 22.7 GiB peak GPU memory\n- RTX PRO 6000 96GB: 8.0s generation, 9.6s editing\n\nA single checkpoint covers text-to-image, multi-image editing, and transparent RGBA output, with native TP/SP inference, LoRA, and OpenAI-compatible APIs.
More from Infra
- Data center debt built for Jane Street sours in secondary trading, yields hit ~11.3% — GaryMarcus · 2026-09-21
- Dev shills autonomous labs: a 4GPU case worth grabbing for local AI — dee_hw · 2026-09-21
- China's CXMT starts mass-producing 11.95nm DRAM, 50% more dies per wafer — mark_k · 2026-09-21
- Dev Slams AI API Billing: No Hard Spend Cap Anywhere, Budget Alerts Fire Too Late — MaverikSh · 2026-09-21
- Huawei name-drops DeepSeek in keynote, but its actual compute share looks like a tiny fraction of Tencent's — teortaxesTex · 2026-09-21
- Custom CUDA shim runs Stable Diffusion on Mac faster than RTX 5090 on Windows, up to 61% quicker — LioDavinchy · 2026-09-21