Alibaba open-sources Qwen-Image-2.1: 7B image model with native RGBA output, claimed to beat closed rivals
thione · x · 2026-09-28
Alibaba's Qwen team released Qwen-Image-2.1, an open-weight image generation and editing model whose visual generation component has just 7B parameters. The team claims it beats most closed models on their own benchmark (independent evals pending), and it runs on consumer GPUs like an RTX 3090.
Key capabilities:
- Native RGBA generation and editing: isolate objects or edit text on transparent layers
- Handles up to 10 reference images at once for group portraits, virtual try-ons, or room design
- Local edits guided by circles, masks, or painted marks
- Architecture changes and KV cache reuse speed up inference, especially with multiple reference images
Available on Hugging Face, GitHub, and Model Scope with a HF demo. The research license bars commercial use; business users must apply for a separate license.
More from Multimodal
- Horse gaits one-shot with Claude Opus 5.5: impressive coat shine and musculature — shekitup · 2026-09-28
- Ad Buyer Ditches Image Models, Uses Seedance 2.5 for Realistic AI UGC Characters — churchkey · 2026-09-28
- One-word Midjourney prompt: 'Pulchritudinous' at --ar 5:4 --v 8.2 — tisch_eins · 2026-09-28
- Prompt Template Makes Qwen Image 2.1 Design Like Closed-Source Models — wjc_5 · 2026-09-28
- A ChatGPT prompt turns GPT Image into an editorial fashion photographer — umesh_ai · 2026-09-28
- Synthesia ships Express-3, its best avatar model, free for users on all plans — synthesiaIO · 2026-09-28