Qwen3.8 27B runs at 50 tok/s with 100k context on 16GB GPU
qaf23 · reddit · 2026-08-29
Author shares a setup to run Qwen3.8-27B with 100k context at 50 tok/s on a 16GB GPU (RTX 4070 Ti SUPER). Key optimizations include using beellama.cpp, asymmetric kvarn5/kvarn4 KV cache quantization to save VRAM, keeping the last 1024 tokens at full precision, and enabling Speculative Decoding (draft-mtp). The setup uses 15.93GB VRAM and achieves near-lossless quality.
Related event: Qwen 3.8 27B Quantization Tests: 200K Context on 16GB VRAM(4 posts)→
More from Infra
- Lightning AI deploys H200s and launches VMs early access — LightningAI · 2026-08-29
- Open Source Personal AI Datacenter with Full Guides — dee_hw · 2026-08-29
- One Beef Burger Equals Lifetime of ChatGPT Use in Water Usage — dc_lawrence · 2026-08-29
- Beginner's Guide to Ollama: Setting Up Local LLMs from Scratch — 4310sy · 2026-08-29
- Dev: 80%+ of market can't afford frontier model tokens — thedealdirector · 2026-08-29
- Hands-on with NVIDIA DLSS 5: Testing Across Multiple Games — ryanshrout · 2026-08-29