Z-Image CyberRealistic v9 GGUF Quant Fits 8GB VRAM, With x_pad_token Fix
Kordeyl · reddit · 2026-10-07
After weeks of work, the author quantized Z-Image Turbo / CyberRealistic v9 to GGUF so it runs on 8GB cards without offloading the text encoder, bundling diffusion model, Qwen3-4B encoder, and VAE in one repo.
Key takeaways:
- Q4K diffusion (5.15GB) + Q4KM encoder (2.50GB) is the 8GB sweet spot
- Quirk: for Z-Image (Lumina2), most mixed S/M/L quant recipes produce byte-identical output to base K — Q4 is the only exception where Q4KS and Q4K genuinely differ
- The xpadtoken/cappadtoken size mismatch error comes from ggml dropping the leading 1 from shapes during quantization; fix by using Z-Image Power Nodes or patching tensor shapes (also fixed in ComfyUI-GGUF PR #392)
- Sampler settings: steps 8 (Turbo distillation — above 10 degrades), cfg 1.0, euler, 1024×1024
- Verified on a GTX 1070 (8GB): 72% VRAM usage with 1GB to spare; example PNGs embed full workflows, drag-and-drop into ComfyUI
More from Infra
- Datology AI open-sources Zephon, a deterministic on-the-fly dataloader that fixes cross-GPU-run inconsistency — josh_wills · 2026-10-07
- Data center water and power fears overblown? Aluminum smelter uses a Boston-sized grid — csuwildcat · 2026-10-07
- Omarchy Linux ships official browser-based remote desktop plugin — juntao · 2026-10-07
- Datology AI open-sources Zephon, a deterministic on-the-fly data loader built for massive ablations — pratyushmaini · 2026-10-07
- Intel CEO confirms continued chip partnership with Musk's Terafab project — elonmusk · 2026-10-07
- Symmetrix-XL: open-source engine simulates 10M atoms on a single GPU — CatAstro_Piyush · 2026-10-07