Qwen 27B at ~18 tok/s on Just 12GB VRAM + 8GB RAM: Full Recipe Released
bodhi371 · reddit · 2026-10-07
A user got Qwen3.8-27B running at 18 tok/s decode and 500-600 tok/s prefill (64k context) on just 12GB VRAM + 8GB RAM using the ISTA-DASLab GSQ-RCO IQ3S quant (11GB) — their best-tested quant, with 0bserverx's Heretic GSQ-RCO variant about 7% slower. The trick: Qwen3.8's hybrid architecture needs KV cache for only 16 of 64 layers, and q40 KV cache shrinks 64k context to 1.1GB vs 4GB at f16. Setup: 9900X + RTX 4070S 12GB, stock llama.cpp with -ngl 58 -ot tokenembd=CPU -ctk q40 -ctv q40 -c 64000. Full scripts and numbers on GitHub (bodhi37/Qwen3.8-27B-12GBVRAM-Recipe).
Related event: Qwen3.8 27B Runs on Consumer GPUs as Low as 12GB VRAM(2 posts)→
More from Infra
- 100B decisions in 5 minutes: a layered agent pipeline where Opus only probes 96 hotspots — Arindam_1729 · 2026-10-07
- OpenSBI and Linux now boot on Maxion cores of ET-SOC1 cards, porting done agentically — glenbeer · 2026-10-07
- Samsung's Q3 profit reportedly set to soar 770% YoY as AI demand overwhelms memory supply — Polymarket · 2026-10-07
- openTPU: A One-Person Open-Source AI Accelerator Designed by AI Agents, Running 10 LLMs on an FPGA — ai · 2026-10-07
- Free Online Conference All Day AI Set for Oct 22 With Talks on SLMs and Local Inference — FikoFox · 2026-10-07
- Best $4,000 Local LLM Rig? Weighing R9700s, Strix Halo, and 6x Arc B60 — BinaryGrind · 2026-10-07