Qwen 27B at ~18 tok/s on Just 12GB VRAM + 8GB RAM: Full Recipe Released

bodhi371 · reddit · 2026-10-07

A user got Qwen3.8-27B running at 18 tok/s decode and 500-600 tok/s prefill (64k context) on just 12GB VRAM + 8GB RAM using the ISTA-DASLab GSQ-RCO IQ3S quant (11GB) — their best-tested quant, with 0bserverx's Heretic GSQ-RCO variant about 7% slower. The trick: Qwen3.8's hybrid architecture needs KV cache for only 16 of 64 layers, and q40 KV cache shrinks 64k context to 1.1GB vs 4GB at f16. Setup: 9900X + RTX 4070S 12GB, stock llama.cpp with -ngl 58 -ot tokenembd=CPU -ctk q40 -ctv q40 -c 64000. Full scripts and numbers on GitHub (bodhi37/Qwen3.8-27B-12GBVRAM-Recipe).

Related event: Qwen3.8 27B Runs on Consumer GPUs as Low as 12GB VRAM(2 posts)→

Original post →

More from Infra

Infra channel →