EXL3 Quantization Test: Running 30B Model on 12GB VRAM Smoothly
PyaesoneP · reddit · 2026-08-30
The author shares experience running Muse Glimmer 30B EXL3-SC 3.00bpw H4 fully resident on a 12GB VRAM GPU at 100K context. It achieves 30 tok/s and shows negligible quality difference compared to the official 17GB K-quant for Hermes Agent use cases. The author also tested Qwen 3.8 27B at SC2.20bpw H3, finding it usable but preferring Unsloth UDQ4KXL for coding tasks.
More from coding & agent
- Agent workflows break in prod despite passing sandbox tests: how to test? — Common_Dream9420 · 2026-08-30
- TrustScoreAgent: Open reputation registry for agent API calls — TrustScoreAgent · 2026-08-30
- BrainAPI beats Mem0/Zep; bottleneck is infra, not models — shbong · 2026-08-30
- Arcturus Bot Integrates GPT-image-2, Showcasing Image Gen in Group Chats — MikePFrank · 2026-08-30
- Adapting ColBERT MaxSim to 10,000-D Bipolar HDC for Edge Memory Gating — Equivalent-Flan-1590 · 2026-08-30
- Varying definitions of Agent memory architecture by tech leaders — bibryam · 2026-08-30