Running 27B model on 12GB VRAM: Qwen 3.8 quantization benchmark

Square_Light1441 · reddit · 2026-08-28

A Reddit user shared a high-efficiency quantization setup for Qwen 3.8 27B, combining QAT Q2 weight quantization with Q5 KV Cache. The total RAM usage is around 13-14GB. Benchmarks show minimal performance degradation, allowing the model to run on a 12GB GPU (at 100K context) with capabilities claimed to surpass Claude Sonnet 4.6.

Related event: Quantized Qwen 3.8 27B Runs 200K Context on Low VRAM(2 posts)→

Original post →

More from Infra

Infra channel →