Qwen 27B Q4 with 100K context at ~30 t/s on a 16GB AMD RX 7800 XT: full guide
Haunting-Stretch8069 · reddit · 2026-09-29
The author runs Qwen 3.8 27B Q4 on a 16GB AMD RX 7800 XT at 30 t/s decode with 100K context, showing it's feasible on 16GB VRAM.
Key steps:
- Build llama.cpp with Vulkan instead of ROCm (ROCm uses more VRAM and leaks);
- Use unsloth's Qwen3.8-27B-UD-IQ4XS.gguf plus mmproj-F16 for vision;
- Quantize KV cache to q80 (K) and q51 (V) to save memory;
- Enable flash-attn, disable unified KV, set 100096 ctx with checkpointing;
- Full copy-paste llama-server command with sampling params included.
More from Infra
- 8B Model Hits 60 Tokens/sec on Phone CPU Alone, No GPU or NPU Needed — const_reborn · 2026-09-29
- Single HBM stack reads Wikipedia 20x per second, HBM5 to double bandwidth by 2028 — lemire · 2026-09-29
- Data centers are more than GPU warehouses: sovereignty means the right to switch systems off — AryHHAry · 2026-09-29
- AsideAI cuts compaction/dreaming token use 7x, doubles cache hit rate — garrytan · 2026-09-29
- Nebius cuts agent training batch collection time by 66.9%, from ~10 min to just over 3 — demian_ai · 2026-09-29
- Celesto: open-source persistent microVM computers for AI agents, boots in 500ms — aniketmaurya · 2026-09-29