Alibaba open-sources Qwen-Image-2.1: 7B image model tops open-source at 60.28
机器之心 · wechat · 2026-09-20
Alibaba's Qwen team open-sourced Qwen-Image-2.1, unifying text-to-image generation and editing in one pipeline aimed at real design workflows.
- Benchmarks: 60.28 on Qwen-Image-Bench, beating NanoBanana2.0 (59.82) and GPT Image 1.5 (59.65) as the top open-source model, with only 7B vision-generation parameters (20-layer Single-Stream DiT, native 2K).
- Capabilities: native RGBA transparent-image generation and subject extraction from photos; up to 10 reference images for multi-image editing with lasso/brush/mask-based local edits; stronger identity and product-packaging fidelity; better text rendering; panorama, infographic, three-view and storyboard generation.
- Efficiency: hybrid-granularity attention — token-level causal masks for text, chunk-level masks for images with KVCache reusing static context — cutting memory and boosting inference on multi-image inputs.
- Weights are on Hugging Face and ModelScope; note the strict non-commercial qwen-research license with no revenue cap.
More from Multimodal
- Ancient phalanx AI video shows a Grok Imagine + Veo + Topaz workflow — creatoroff · 2026-09-20
- Comfy-Org's single-file Qwen-Image-2.1 model trends on Hugging Face — Comfy-Org · 2026-09-20
- Qwen Image 2.1 regeneration removes the noisy texture in GPT images — Choidonhyeon · 2026-09-20
- Open image-to-3D on 6GB VRAM: Trellis 2, Hunyuan3D and more all fall short for characters — danihyder47 · 2026-09-20
- Open-Weight 7B Image Model Claimed to Beat Nano Banana 2.0, Unverified — QuixiAI · 2026-09-20
- NoSpoon music video agent in closed beta: full MV in minutes from one track — Kyrannio · 2026-09-20