Qwen 27B with vision on a 16GB GPU: 85K context at 45 tok/s, full config
FerLuisxd · reddit · 2026-09-10
A detailed local-deployment recipe for running Qwen 3.8 27B with vision on a 5060 Ti (16GB VRAM): IQ3XXS-mtp GGUF quant from ISTA-DASLab, BF16 mmproj, via beellama.cpp. Result: 45 tok/s decode, 300 tok/s prefill, 85K context, with 1.5GB headroom left.
Key tricks: full GPU offload with direct-io, kvarn4 KV-cache compression for both K and V (author links benchmarks showing acceptable quality loss), draft-MTP speculative decoding (max 2 draft tokens), and the option to move mmproj to CPU to free more VRAM. Full config block included; author invites other 16GB sweet-spot configs.
More from Infra
- colibri: Pure-C Zero-Dependency Engine Streams MoE Experts From Disk to Run Frontier Models Locally — JustVugg · 2026-09-10
- HBM seen topping 40% revenue share next year as Samsung challenges SK Hynix's lead — zephyr_z9 · 2026-09-10
- A quiet CPU shortage is hitting cloud users as hyperscalers shrug, devs report — mgill25 · 2026-09-10
- "Going for infinite energy": the Kardashev II mindset meme in AI circles — MarvinTBaumann · 2026-09-10
- Redis LangCache claims 90% LLM cost cuts; KV, prefix, prompt and semantic caching explained — blaizedsouza · 2026-09-10
- Raspbian creator laments $300 Raspberry Pi as AI demand prices out young programmers — evilsocket · 2026-09-10