Quantized Krea-2-Turbo Runs on 6GB VRAM, Humming Kernel Hits 1.6x Speedup

ali_byteshape · reddit · 2026-08-12

To address high VRAM costs, developers have released quantized versions of the Krea-2-Turbo image generation model, available in GGUF format for ComfyUI and Humming format for vLLM-Omni.

The Humming path is currently experimental and limited to Linux + NVIDIA.

Original post →

More from Infra

Infra channel →