Running Krea 2 Turbo on 8GB VRAM: RTX 3070 Ti Local Test
niechta · reddit · 2026-08-13
A developer shared a working configuration for running the Krea 2 Turbo model on a single RTX 3070 Ti (8GB VRAM).
- Quantization & Optimization: The model was quantized to 4-bit (W4A4), reducing its size to 6.9 GB, with LoKr weights baked into the bf16 base weights prior to quantization.
- Software Stack: Runs on ComfyUI 0.32, utilizing INT8 attention optimizations and a Qwen3-VL-4B text encoder in fp8.
- Performance: Generating a 1024×1024 image takes about 11 seconds at 8 steps.
More from Infra
- Over-provisioning AI Agents? Mixing Model Tiers Cuts 75% of Costs — AIMOWAY · 2026-08-13
- Crypto Market Maker Wintermute to Invest ~$1B in AI Data Centers — Polymarket · 2026-08-13
- Nvidia Partners with Wall Street to Mobilize $500B for AI Infrastructure — fortune · 2026-08-13
- New Quantization Framework to Run 1.6T Models on a Single B300 GPU at 50 tok/s — dosco · 2026-08-13
- Menlo Park Daytime Electricity Hits 53.8¢/kWh: Running Own GPUs Becomes Irrational — generativist · 2026-08-13
- Fluidstack Visits NYSE to Discuss US AI Infrastructure Investment — MxMnr · 2026-08-13