Running 27B model on 12GB VRAM: Qwen 3.8 quantization benchmark
Square_Light1441 · reddit · 2026-08-28
A Reddit user shared a high-efficiency quantization setup for Qwen 3.8 27B, combining QAT Q2 weight quantization with Q5 KV Cache. The total RAM usage is around 13-14GB. Benchmarks show minimal performance degradation, allowing the model to run on a 12GB GPU (at 100K context) with capabilities claimed to surpass Claude Sonnet 4.6.
Related event: Quantized Qwen 3.8 27B Runs 200K Context on Low VRAM(2 posts)→
More from Infra
- Autonomous launches $26,100 dual-RTX 5090 AI workstation, 'The Diablo' edition — dee_hw · 2026-08-28
- oMLX 0.6.3 Released: Qwen/GLM Flash Support, 160% Faster Decode — HankYeomans · 2026-08-28
- vLLM benchmarks MTP, EAGLE-3, and other speculative decoding methods on AMD GPUs — vllm_project · 2026-08-28
- Data center revenue drives $79M rate cut plan by US utility I&M — toptickcrypto · 2026-08-28
- Why AI data centers still rely on water cooling over air cooling? — wren337 · 2026-08-28
- Microsoft Tutorial: Configure AI Gateway in Foundry Resources — adnan_hashmi · 2026-08-28