Running Muse Glimmer 30B with 256k Context on a Single RTX 3090: Benchmarks

coder543 · reddit · 2026-08-10

A developer tested the local deployment of the Muse Glimmer 30B model on a single RTX 3090. Using Q4KXL quantization, DFlash, and mmproj, the model comfortably supports a 256k context window while consuming only about 22-23GB of VRAM, significantly outperforming similar-sized models like Qwen3.6-27B and Gemma-4-31B.

Performance:

In contrast, Qwen3.6-27B can only handle around 125k tokens with Q8 KV cache on the same GPU. Muse Glimmer's efficient VRAM usage eliminates the need for more expensive hardware like the DGX Spark for running full context.

Original post →

More from Infra

Infra channel →