Deploying 30B Models on RTX 3090: Balancing VRAM and Quantization
ydnar · reddit · 2026-08-11
A developer shares their configuration and optimization experience running the Muse Glimmer 30B model on an RTX 3090 (24GB) using llama-server.
Configuration
- Model: unsloth/Muse-Glimmer-30B-GGUF:UD-Q4KXL
- Context: 131072
- Cache: K/V cache set to q80 with Flash Attention enabled.
- Generation: temp 1.0, top-k 64, top-p 0.95.
Status & Discussion
Currently, nvidia-smi shows 18163MiB / 24575MiB VRAM usage, leaving some headroom. The author found it performed better than Qwen 27B in coding/logic tasks.
The author asks the community whether to squeeze in a higher precision quantization (UD-Q5KXL) for better quality, or stick with Q4 to maximize tokens/s throughput.
More from Infra
- Oz-FP4: Emulating FP64 DGEMM on Low-Precision FP4 Tensor Cores — teortaxesTex · 2026-08-11
- 5x Speedup for Local Video Generation: WanGP Optimizes Wan2.1 — cocktailpeanut · 2026-08-11
- Training an EAGLE-3 Speculative Decoding Drafter for Gemma-3-27B on a Single RTX 5090 — max_paperclips · 2026-08-11
- Future AI Compute: Free Energy and Kimi K5 to Unlock a $5T Market — MarvinTBaumann · 2026-08-11
- Testing 16 Quantization Schemes for Qwen 27B: GGUF Offers Best Quality-Size Tradeoff — Hefty_Wolverine_553 · 2026-08-11
- Google Cloud Revenue Jumps 82%, $514B Backlog Validates AI Demand — DavidLinthicum · 2026-08-11