Qwen-Image-2.1 model card: mixed-granularity attention, prefix KV reuse, quickstart code
AdinaYakup · x · 2026-09-21
The Hugging Face model card details Qwen-Image-2.1: a 7B-parameter (32-layer Single-Stream DiT) model using mixed-granularity attention and prefix KV cache reuse for low-cost inference. It unifies generation and editing — RGBA transparent image generation, editing transparent layers, subject extraction — supports up to 10 reference images with identity preservation, and enables local edits via circles, painted annotations or separate masks. Typography, portrait lighting and fine textures are improved. The card ships pip install instructions and a Diffusers QwenImage21Pipeline quickstart; the license is research-and-evaluation-only.
More from Multimodal
- Qwen-Image-2.1: 9 simple ComfyUI workflows, near pixel-perfect editing — nomadoor · 2026-09-21
- Turn any video into volumetric playback with exportable 360° orbits at custom tempo — LinusEkenstam · 2026-09-21
- One image plus one prompt generates an entire cinematic getaway video — umesh_ai · 2026-09-21
- Full Krea2 concept-art pipeline: four custom LoRAs, 5MP native frames, two-step upscale to 7000px — Dacrikka · 2026-09-21
- Phone video in, navigable 3D room out: K3-powered splat pipeline runs in the browser — willeastcott · 2026-09-21
- Viral Arabic AI trick: upload a photo and your name to get a magazine-cover portrait — aziz4ai · 2026-09-21