Gemma 4 32B Causes Frequent OOM Crashes on RTX 4090
BSPiotr · reddit · 2026-08-05
A user reported frequent Out of Memory (OOM) and CUDA errors when running the Gemma 4 32B model (Q4KM, 32k context) locally on an RTX 4090. Even with SWA enabled and theoretical VRAM usage fitting within 24GB, inference tools like Kobold and Ooba still crash. The issue appears isolated to the Gemma 4 family, prompting the user to ask if it's an NVIDIA config issue, an SWA bug, or a problem with llama.cpp.
More from Infra
- Ex-OpenAI Exec Slams Goldman Sachs Token Demand Forecast, Cites 100x Cost Drop — ChrSzegedy · 2026-08-05
- SK Hynix and Samsung Evaluate AMEC Etchers for Chinese Fabs — zephyr_z9 · 2026-08-05
- NVIDIA Open-Sources CuTe Algebra and Compiler Stack to Boost AI Kernel Agents — GregoryDiamos · 2026-08-05
- Struggles of Running Local TTS Inference on AMD GPUs in Windows — aboutthednm · 2026-08-05
- Running 1.5B Voice Model Locally on iPhone: Only 2.2GB Memory — Acceptable-Cycle4645 · 2026-08-05
- Influencer Rejects AI Hype Claims: Intelligence Will Soon Drive the Physical World — DeryaTR_ · 2026-08-05