Diffusers adds Nunchaku 4-bit diffusion inference with up to 50% less VRAM

dl_weekly · x · 2026-07-28

Hugging Face says Diffusers now natively supports Nunchaku’s SVDQuant W4A4 kernels, making 4-bit diffusion inference available through a simple frompretrained() load.

The integration aims to make modern image-generation models usable on consumer GPUs by cutting peak VRAM by up to 50% and improving speed by 30–80%. The post also notes that the companion diffuse-compressor toolkit lets users quantize new architectures and publish them as standard Diffusers repos.

Original post →

More from Infra

Infra channel →