Running Muse Glimmer 30B with 256k Context on a Single RTX 3090: Benchmarks
coder543 · reddit · 2026-08-10
A developer tested the local deployment of the Muse Glimmer 30B model on a single RTX 3090. Using Q4KXL quantization, DFlash, and mmproj, the model comfortably supports a 256k context window while consuming only about 22-23GB of VRAM, significantly outperforming similar-sized models like Qwen3.6-27B and Gemma-4-31B.
Performance:
- Inference speed ranges from 64 to 124 tok/s depending on the content.
- Prompt processing reaches up to 1400 tok/s.
In contrast, Qwen3.6-27B can only handle around 125k tokens with Q8 KV cache on the same GPU. Muse Glimmer's efficient VRAM usage eliminates the need for more expensive hardware like the DGX Spark for running full context.
More from Infra
- AI Cybersecurity Boom: Defending Against AI Attacks Requires More AI — jzl86 · 2026-08-10
- Red Hat AI Releases Muse-Glimmer 30B FP8 Quantized Checkpoint, Halving Memory — vllm_project · 2026-08-10
- Microsoft Plans Maia 300 AI Chip Ramp, Targeting Over 1M Units by 2027 — thoefler · 2026-08-10
- OpenAI Hires Power Trading Lead to Hedge Energy for Data Centers — nathanbenaich · 2026-08-10
- Developer Achieves 114 TPS on 30B Model Using Quantization and Slicing — dejavucoder · 2026-08-10
- Help Wanted: Running Minimax H3 on an 8GB VRAM RTX 2070 Super? — rdwulfe · 2026-08-10