Qwen3.8 27B runs at 50 tok/s with 100k context on 16GB GPU

qaf23 · reddit · 2026-08-29

Author shares a setup to run Qwen3.8-27B with 100k context at 50 tok/s on a 16GB GPU (RTX 4070 Ti SUPER). Key optimizations include using beellama.cpp, asymmetric kvarn5/kvarn4 KV cache quantization to save VRAM, keeping the last 1024 tokens at full precision, and enabling Speculative Decoding (draft-mtp). The setup uses 15.93GB VRAM and achieves near-lossless quality.

Related event: Qwen 3.8 27B Quantization Tests: 200K Context on 16GB VRAM(4 posts)→

Original post →

More from Infra

Infra channel →