Z Image HSWQ Hybrid Quantization Boosts VRAM and Speed

Zestyclose_Bake3680 · reddit · 2026-08-20

A developer released a Hybrid-Sensitivity-Weighted Quantization (HSWQ) method and ComfyUI loader for the Z Image model. Based on ConvRot INT8, this technique uses a backward sweep to downgrade non-essential layers to 4-bit (NVFP4) while maintaining critical layers in INT8. Benchmarks show Z Image has significantly higher quantization robustness than SDXL and Krea2, achieving SSIM scores over 0.99. While file sizes remain similar to INT8, VRAM usage and generation speed are significantly improved. The release includes technical guides documenting the shift from previous HSWQ theories towards a 'discarding' philosophy and the discovery of inter-layer dependencies.

Related event: Z Image Introduces HSWQ Hybrid Quantization for Lower VRAM Usage(2 posts)→

Original post →

More from Infra

Infra channel →