New Q8_CR GGUF format keeps Krea 2 diffusion speed near INT8 while shrinking VRAM pressure
molbal · reddit · 2026-07-27
What changed
- A Reddit developer introduced Q8CR, a new GGUF format for diffusion models that combines pre-rotated INT8 weights with FP32 per-row scales and keeps small or precision-sensitive tensors in higher precision.
- The format is designed first for Ampere and Turing GPUs and takes advantage of ComfyUI's native INT8 ConvRot path when available.
Reported results
- On Krea 2, the author says Q8CR lands very close to the INT8 baseline speed while staying smaller than the full INT8 file.
- Measured sizes/speeds:
- INT8: 11.95 GB, baseline speed
- Q80: 13.56 GB, about 60.2% of baseline speed
- Q8CR: 12.84 GB, about 100.4% of baseline speed
Why GGUF
- The author argues GGUF is better than safetensors for low-VRAM inference because weights can stay quantized in system RAM and be streamed layer by layer into VRAM.
- Full INT8 safetensors still tend to load into normal FP16/BF16 execution paths, which can increase transfer and memory pressure.
Availability
- A GGUF file for Krea 2 Turbo Q8CR and a ComfyUI-GGUF custom node are linked.
- The author says they will test the format on Flux 2 Klein, Z-Image, and Ideogram next.
More from Infra
- AI-Trader adds an MCP server so LLMs can run trading backtests — tom_doerr · 2026-07-27
- Nvidia signs $1.5B multi-year Amkor deal to expand US chip packaging capacity — Beth_Kindig · 2026-07-27
- PyTorch DDP misses a tiny-parameter NVIDIA GPU optimization out of the box — gordic_aleksa · 2026-07-27
- A fully local AI girlfriend runs on 15GB VRAM with Whisper, llama.cpp, and Qwen3-TTS — max_paperclips · 2026-07-27
- OpenAI reportedly plans to spend over $30B on a 3.2GW Georgia data center — Beth_Kindig · 2026-07-27
- Two RTX 5080s failed to beat one RTX 5090 in ComfyUI single-render tests — Geekdomo · 2026-07-27