Krea2 Quantization Formats Compared
y3kdhmbdb2ch2fc6vpm2 · reddit · 2026-07-13
A comparison of Krea2 across various quantization/precision formats: BF16, INT4convrot, INT8convrot, GGUFQ8, FP8scaled, and NVFP4.
Experiment Setup
- Generated 126 images, including scenarios with LoRA
- Same prompt and seed used for each comparison group
- Test environment: RTX 5070 Ti 16GB + 32GB DDR5 + NVMe
- ComfyUI 0.27.0, PyTorch 2.12.0+cu130, default attention
- Resolution 1024×1024, Euler / simple, 8 steps, cfg 1.0, wan 2.1 fp32 VAE, qwen 3vl 4b bf16 clip
Key Findings
- Second-generation speeds were generally faster, but the overall ranking remained consistent
- Generation times were roughly:
- BF16: 22.00 s → 13.36 s
- GGUFQ8: 41.46 s → 35.82 s
- INT8convrot: 8.13 s → 5.88 s
- FP8scaled: 10.77 s → 9.35 s
- NVFP4: 9.33 s → 7.78 s
- INT4convrot: 7.15 s → 5.84 s
- Using one or more LoRAs did not affect generation time
- In this test, INT4convrot was the fastest, while GGUFQ8 was the slowest
Related event: Benchmarking Krea 2 Turbo Across Quantization Formats(3 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22