Alibaba open-sources Qwen-Image-2.1: 7B unified text-to-image and editing model with RGBA support

Xianbao_QIAN · x · 2026-09-20

Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a compact 7B visual generation component (32 Single-Stream DiT layers), under the qwen-research license, ready for vLLM-Omni and Diffusers.

Four key improvements:

Available on Hugging Face and ModelScope; quick start via the QwenImage21Pipeline in diffusers (bfloat16, native 2048×2048).

Related event: Alibaba Open-Sources Qwen-Image-2.1, a Unified 7B Image Generation and Editing Model(20 posts)→

Original post →

More from Multimodal

Multimodal channel →