Alibaba open-sources Qwen-Image-2.1: 7B image gen/edit model with RGBA and 10-image references
multimodalart · x · 2026-09-20
Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and editing model with a compact 7B visual component (32 Single-Stream DiT layers). Key upgrades: mixed-granularity attention with prefix KV cache reuse for efficiency; native RGBA transparent image generation and layer editing; versatile editing with up to 10 reference images plus circle/annotation/mask-based local edits with identity preservation; and improved typography, portrait lighting and textures. Available in Diffusers, supporting 2048×2048 output under a qwen-research license.
More from Multimodal
- Comfy-Org's single-file Qwen-Image-2.1 model trends on Hugging Face — Comfy-Org · 2026-09-20
- Qwen Image 2.1 regeneration removes the noisy texture in GPT images — Choidonhyeon · 2026-09-20
- Open image-to-3D on 6GB VRAM: Trellis 2, Hunyuan3D and more all fall short for characters — danihyder47 · 2026-09-20
- Open-Weight 7B Image Model Claimed to Beat Nano Banana 2.0, Unverified — QuixiAI · 2026-09-20
- NoSpoon music video agent in closed beta: full MV in minutes from one track — Kyrannio · 2026-09-20
- AI Agent Beepboop Releases First Self-Written, Self-Directed Music Video — Kyrannio · 2026-09-20