Qwen-Image 2.1 open weights land with day-0 SGLang support, runs on one RTX 4090
ying11231 · x · 2026-09-21
Alibaba's Qwen team released open-weight Qwen-Image-2.1, a 7B unified generation/editing model with native RGBA transparency and up to 10 reference images. SGLang shipped day-0 support: native precision on a single RTX 4090 24GB with CPU offload (18.7s for 1024×1024, 21.7s editing, 22.7 GiB peak), 8.0s/9.6s on RTX PRO 6000, with TP/SP, LoRA and OpenAI-compatible APIs at 40 denoising steps, no quantization.
More from Infra
- Tobi Lütke: local Dell server runs DeepSeek 4.1 Flash at ~300 tok/s, a billion tokens a month — BLUECOW009 · 2026-09-21
- Running Qwen3.8-27B EXL3 on RTX 3060 + 5060 Ti: 50 tok/s with tensor parallelism and MTP — bring_back_the_v10s · 2026-09-21
- Baseten CEO says token volume grew 40x YoY while revenue grew ~10x in 12 months — rohanpaul_ai · 2026-09-21
- AI doesn't live in the cloud: who pays the environmental price of scale? — SuzannahB1001 · 2026-09-21
- AI Cluster Bottleneck Isn't Chips — It's the Lasers Moving Data Between Them — McDonaghMatthew · 2026-09-21
- Inside SemiAnalysis: the research firm guiding a $1T AI infrastructure buildout — AccBalanced · 2026-09-21