Qwen open-sources Qwen-Image-2.1: 7B unified generation and editing with RGBA
linoy_tsaban · x · 2026-09-20
Alibaba's Qwen team open-sourced Qwen-Image-2.1, a unified text-to-image generation and editing model with a compact 7B visual stack (32 single-stream DiT layers). Key upgrades: mixed-granularity attention plus prefix KV cache reuse for efficiency; native RGBA transparency (generate transparent images, edit layers, extract subjects); editing with up to 10 reference images and circle/paint/mask local edits with identity preservation; improved typography, portrait lighting and textures. Weights, Space and blog are live on Hugging Face, with one-line Diffusers loading.
More from Multimodal
- Wiring Codex Astra to Hyper3D Rodin MCP turns prompts into clickable 3D prototypes — alex_verem · 2026-09-20
- NoSpoon Agent Autonomously Generates Microdramas That Hook Its Own Developer — Kyrannio · 2026-09-20
- Creator generates 10 creepy SCP-096 images with AI, eyes a short film — VraserX · 2026-09-20
- First hour with Qwen Image 2.1: editing hit-or-miss, 2511 still wins some comparisons — LowYak7176 · 2026-09-20
- Qwen-Image-2.1 ships under strict non-commercial research license, no revenue cap — ostrisai · 2026-09-20
- Full 1961 TV drama generated with Deedance 2.5 on Runway, period details praised — azed_ai · 2026-09-20