Krea 2 Turbo Quantization Speed Benchmarks
Merserk13 · reddit · 2026-07-13
Using an identical ComfyUI generation pipeline on an NVIDIA RTX PRO 6000 Blackwell, the author conducted speed and VRAM benchmarks across 8 model formats for Krea 2 Turbo.
Key Findings
- 1024×1024: int8convrot is the fastest at a 2.424s median; int4convrot is only 154ms slower but features lower peak VRAM usage.
- 2048×2048: int4convrot overtakes int8convrot, taking the crown with a 12.678s median.
- Native quantized checkpoints generally outperform GGUF versions on this Blackwell machine. Notably, GGUF Q4KM does not run faster despite having a smaller file size.
Resource Usage Observations
- int4convrot offers the most balanced profile: near-top speed combined with some of the lowest VRAM consumption.
- nvfp4 also demonstrates strong performance, with speeds approaching int4 and similarly low VRAM overhead.
- bf16 demands the highest VRAM while lagging noticeably behind quantized versions in speed.
Related event: Benchmarking Krea 2 Turbo Across Quantization Formats(3 posts)→
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22