Quantized Krea-2-Turbo Runs on 6GB VRAM, Humming Kernel Hits 1.6x Speedup
ali_byteshape · reddit · 2026-08-12
To address high VRAM costs, developers have released quantized versions of the Krea-2-Turbo image generation model, available in GGUF format for ComfyUI and Humming format for vLLM-Omni.
- VRAM Requirements: Scaled down from 6.14 GB (3.83 bpw) up to 14.31 GB (8.93 bpw).
- Performance: On an RTX 5090 at 1024×1024, Humming kernels run about 1.6x faster per step than GGUF (0.41s vs. 0.67s), taking just 4.0s for an 8-step generation.
- Image Quality: The team tested all quants against the BF16 baseline across 24 prompts and provided an interactive online slider for side-by-side comparison.
The Humming path is currently experimental and limited to Linux + NVIDIA.
More from Infra
- Rethinking AI Energy Metrics: Joules per Token Must Account for Quality and Task — prateekj · 2026-08-13
- Can MiniMax H3 Run on a Laptop with Only 32GB of RAM? — Puzzleheaded_Emu8419 · 2026-08-13
- Bittensor Subnet Runs Full Kimi K3 on 80 RTX 5090s, Halving API Costs — markjeffrey · 2026-08-13
- US DOE to Push 20% of Budget into Compute Infrastructure for AI — thoefler · 2026-08-13
- CoreWeave and W&B to Host Fully Connected Conference on Production AI Infrastructure — wandb · 2026-08-13
- vLLM Announces Day-0 Support for Qwen3.8 2.4T with Ready-to-use 4-bit Checkpoints — vllm_project · 2026-08-13