Qwen-Image 2.1 Merged into diffusers with Block-Causal Attention and 64-Channel VAE

linoy_tsaban · x · 2026-09-18

Qwen-Image 2.1, a unified text-to-image and image-to-image model billed as the best value-for-compute in the Qwen-Image family, has been merged into Hugging Face diffusers (PR #14804). It features a single-stream transformer with block-causal attention and KV cache, a 64-channel VAE, and a pipeline supporting both text-to-image and image-conditioned generation.

Original post →

More from Multimodal

Multimodal channel →