Alibaba open-sources Qwen-Image-2.1: 7B unified text-to-image and editing model with RGBA support
Xianbao_QIAN · x · 2026-09-20
Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a compact 7B visual generation component (32 Single-Stream DiT layers), under the qwen-research license, ready for vLLM-Omni and Diffusers.
Four key improvements:
- Compact & efficient: mixed-granularity attention plus prefix KV cache reuse keeps quality at low compute
- Native transparency: generate/edit RGBA transparent layers and extract subjects from photos in one model
- Versatile editing: up to 10 reference images, local edits via circles, painted annotations or masks, with identity preservation for people and products
- Better realism: improved typography, portrait lighting and fine details
Available on Hugging Face and ModelScope; quick start via the QwenImage21Pipeline in diffusers (bfloat16, native 2048×2048).
More from Multimodal
- Ancient phalanx AI video shows a Grok Imagine + Veo + Topaz workflow — creatoroff · 2026-09-20
- Comfy-Org's single-file Qwen-Image-2.1 model trends on Hugging Face — Comfy-Org · 2026-09-20
- Qwen Image 2.1 regeneration removes the noisy texture in GPT images — Choidonhyeon · 2026-09-20
- Open image-to-3D on 6GB VRAM: Trellis 2, Hunyuan3D and more all fall short for characters — danihyder47 · 2026-09-20
- Open-Weight 7B Image Model Claimed to Beat Nano Banana 2.0, Unverified — QuixiAI · 2026-09-20
- NoSpoon music video agent in closed beta: full MV in minutes from one track — Kyrannio · 2026-09-20