EXL3 Quantization Test: Running 30B Model on 12GB VRAM Smoothly

PyaesoneP · reddit · 2026-08-30

The author shares experience running Muse Glimmer 30B EXL3-SC 3.00bpw H4 fully resident on a 12GB VRAM GPU at 100K context. It achieves 30 tok/s and shows negligible quality difference compared to the official 17GB K-quant for Hermes Agent use cases. The author also tested Qwen 3.8 27B at SC2.20bpw H3, finding it usable but preferring Unsloth UDQ4KXL for coding tasks.

Original post →

More from coding & agent

coding & agent channel →