Alibaba open-sources Qwen-Image-2.1: 7B image gen/edit model with RGBA and 10-image references

multimodalart · x · 2026-09-20

Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and editing model with a compact 7B visual component (32 Single-Stream DiT layers). Key upgrades: mixed-granularity attention with prefix KV cache reuse for efficiency; native RGBA transparent image generation and layer editing; versatile editing with up to 10 reference images plus circle/annotation/mask-based local edits with identity preservation; and improved typography, portrait lighting and textures. Available in Diffusers, supporting 2048×2048 output under a qwen-research license.

Related event: Alibaba Open-Sources Qwen-Image-2.1, a Unified 7B Image Generation and Editing Model(20 posts)→

Original post →

More from Multimodal

Multimodal channel →