Alibaba open-sources Qwen-Image-2.1: 7B unified image generation and editing model with native RGBA support
RisingSayak · x · 2026-09-20
Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and editing model, with Day-0 support in Diffusers.
Key points:
- Only 7B parameters (32 Single-Stream DiT layers) with mixed-granularity attention and prefix KV cache reuse for efficient inference
- Natively generates and edits RGBA transparent images and extracts subjects from photos
- Supports up to 10 reference images, local edits via circles, painted annotations, or masks, with identity preservation for people and products
- Improved typography, portrait lighting, and fine details
The model is available on Hugging Face as Qwen/Qwen-Image-2.1 with a quick-start pipeline.
More from Multimodal
- Comfy-Org's single-file Qwen-Image-2.1 model trends on Hugging Face — Comfy-Org · 2026-09-20
- Qwen Image 2.1 regeneration removes the noisy texture in GPT images — Choidonhyeon · 2026-09-20
- Open image-to-3D on 6GB VRAM: Trellis 2, Hunyuan3D and more all fall short for characters — danihyder47 · 2026-09-20
- Open-Weight 7B Image Model Claimed to Beat Nano Banana 2.0, Unverified — QuixiAI · 2026-09-20
- NoSpoon music video agent in closed beta: full MV in minutes from one track — Kyrannio · 2026-09-20
- AI Agent Beepboop Releases First Self-Written, Self-Directed Music Video — Kyrannio · 2026-09-20