Deploying 30B Models on RTX 3090: Balancing VRAM and Quantization

ydnar · reddit · 2026-08-11

A developer shares their configuration and optimization experience running the Muse Glimmer 30B model on an RTX 3090 (24GB) using llama-server.

Configuration

Status & Discussion

Currently, nvidia-smi shows 18163MiB / 24575MiB VRAM usage, leaving some headroom. The author found it performed better than Qwen 27B in coding/logic tasks.

The author asks the community whether to squeeze in a higher precision quantization (UD-Q5KXL) for better quality, or stick with Q4 to maximize tokens/s throughput.

Related event: Extreme Local Inference: Single GPUs Run 30B Models with Massive Context and High TPS(8 posts)→

Original post →

More from Infra

Infra channel →