Running Qwen3.8 27B on 16GB VRAM: IQ3 quant hits 131k context at 9-30 tps
randomgenericbot · reddit · 2026-10-07
A detailed hands-on guide to running Qwen3.8 27B on a 16GB GPU (5060Ti, single-channel DDR5). The short answer: it works, with caveats.
Key points:
- Use the GSQ-RCO IQ3S quant from ISTA-DASLab; the author benchmarked it side-by-side against a hosted Q8 on an H200 and couldn't tell them apart except for speed
- For large context you must build a custom llama.cpp fork with adaptive KV streaming; this forfeits MTP, which only added 5tps anyway
- Measured: 30tps at empty context, dropping to 9-10tps near 131k context; prefill declines steadily (880tps at 14k tokens, 630tps at 94k)
- Tuning: a 1.5G KV window is a sweet spot (30k context fully in VRAM, headroom for other GPU tasks); Q8 KV cache gets 32k-68k context on stock llama
- Practical path: load the model on stock llama, use a harness like pi.dev to have the model itself help compile the fork — copy-pasting console output in chat mode is enough, no kernel-hacking required
More from Infra
- PyTorch Replaces CUDA with FBTriton for Embedding Kernels: 1.28x Faster Forward, 2x Backward — PyTorch · 2026-10-07
- Nvidia B200 Prices Are Skyrocketing Amid Intense AI Compute Scramble — matt_slotnick · 2026-10-07
- Google & MIT's Coco: an agent platform for TPU hardware-model co-design — dair_ai · 2026-10-07
- Dev builds Rust desktop API gateway unifying OpenAI/Anthropic APIs on one GPU — rootshelldev · 2026-10-07
- PyTorch Conference to showcase DeepSpeed's tensor, sequence, and expert parallelism beyond ZeRO — PyTorch · 2026-10-07
- SpaceX reportedly seeks $40B to fund Nvidia AI chip purchases, Apollo leading — Polymarket · 2026-10-07