Diffusers adds Nunchaku 4-bit diffusion inference with up to 50% less VRAM
dl_weekly · x · 2026-07-28
Hugging Face says Diffusers now natively supports Nunchaku’s SVDQuant W4A4 kernels, making 4-bit diffusion inference available through a simple frompretrained() load.
The integration aims to make modern image-generation models usable on consumer GPUs by cutting peak VRAM by up to 50% and improving speed by 30–80%. The post also notes that the companion diffuse-compressor toolkit lets users quantize new architectures and publish them as standard Diffusers repos.
More from Infra
- Amazon Makes Largest-Ever Donation to Lean FRO for Mathematically Proving AI Agent Behavior — ChrSzegedy · 2026-07-28
- LangChain says open source is in its blood as it partners with NVIDIA — Hacubu · 2026-07-28
- Morgan Stanley maps the AI infrastructure stack and who gets paid at each layer — SumitGup · 2026-07-28
- Google Cloud adds spend caps and sandboxes for untrusted agent workloads in Cloud Run — steren · 2026-07-28
- A Fields Medalist joins OpenAI as $250B data-center financing and Kimi K3 land — Dapper-Tale-4021 · 2026-07-28
- Bloomberg map shows Nvidia at the center of circular AI financing — SumitGup · 2026-07-28