Qwen-Image 2.1 Merged into diffusers with Block-Causal Attention and 64-Channel VAE
linoy_tsaban · x · 2026-09-18
Qwen-Image 2.1, a unified text-to-image and image-to-image model billed as the best value-for-compute in the Qwen-Image family, has been merged into Hugging Face diffusers (PR #14804). It features a single-stream transformer with block-causal attention and KV cache, a 64-channel VAE, and a pipeline supporting both text-to-image and image-conditioned generation.
More from Multimodal
- H3 Animate Tested: Best-in-Class Video Editing Without Extra Pose Estimation — linoy_tsaban · 2026-09-18
- Behind He Tongxue's iPhone 18 Pro review: AI woven into the motion design workflow — tinyfool · 2026-09-18
- Seedance 2.5 generates cinematic Japanese classroom zombie scene, full prompt shared — umesh_ai · 2026-09-18
- TripoAI P2.0 ships Smart UV: AI unwrap in 5-7 seconds per mesh, game-ready chars in hours — majidmanzarpour · 2026-09-18
- Motion designer of 10 years returns to test GPT-6 Astra as a studio director, shares 5-step workflow — PrajwalTomar_ · 2026-09-18
- Why Midjourney still stands apart: creative artistry over strict prompt adherence — umesh_ai · 2026-09-18