Running Krea 2 on 8GB VRAM Causes RAM Leaks; GGUF Becomes the Only Stable Workaround
Full-Belt3640 · reddit · 2026-08-13
A developer using an 8GB AMD RDNA2 GPU and 32GB RAM on Linux shared their struggles with running the Krea 2 model. When using fp8 or int8 quantizations, RAM usage spikes to 99% after the first generation, freezing the system entirely, even if cache clearing is attempted.
Currently, GGUF quants around 7-8GB in size are the only way to run the model reliably, though very few Krea 2 models offer GGUF versions. The author notes that while ComfyUI's dynamic memory management was supposed to make GGUFs obsolete, unsupported hardware like RDNA2 cards still face significant hurdles.
More from Infra
- Meta MTIA Architecture Revealed: 72x MTIA 400 Scale-Up and AEC Cable Design — jwt0625 · 2026-08-13
- Pure Rust Browser Port: Qwen3-TTS Runs Locally Without GPU — doodlestein · 2026-08-13
- Anthropic Reportedly in Talks to Acquire AI Chip Efficiency Startup Decart for ~$6B — coinfanking · 2026-08-13
- Meta MTIA 300 Architecture: Why Only 16 Scale-Up Domain With 96 SerDes Lanes? — jwt0625 · 2026-08-13
- Pure Rust Open-Source Tool Enables Local Multi-Speaker ASR & Diarization in Browser — doodlestein · 2026-08-13
- M3 Ultra Hits 374 tok/s Decoding Running LiquidAI 3B Locally at 16 Concurrent Requests — JosephJacks_ · 2026-08-13