Gemma 4 32B Causes Frequent OOM Crashes on RTX 4090

BSPiotr · reddit · 2026-08-05

A user reported frequent Out of Memory (OOM) and CUDA errors when running the Gemma 4 32B model (Q4KM, 32k context) locally on an RTX 4090. Even with SWA enabled and theoretical VRAM usage fitting within 24GB, inference tools like Kobold and Ooba still crash. The issue appears isolated to the Gemma 4 family, prompting the user to ask if it's an NVIDIA config issue, an SWA bug, or a problem with llama.cpp.

Original post →

More from Infra

Infra channel →