Qwen 27B with vision on a 16GB GPU: 85K context at 45 tok/s, full config

FerLuisxd · reddit · 2026-09-10

A detailed local-deployment recipe for running Qwen 3.8 27B with vision on a 5060 Ti (16GB VRAM): IQ3XXS-mtp GGUF quant from ISTA-DASLab, BF16 mmproj, via beellama.cpp. Result: 45 tok/s decode, 300 tok/s prefill, 85K context, with 1.5GB headroom left.

Key tricks: full GPU offload with direct-io, kvarn4 KV-cache compression for both K and V (author links benchmarks showing acceptable quality loss), draft-MTP speculative decoding (max 2 draft tokens), and the option to move mmproj to CPU to free more VRAM. Full config block included; author invites other 16GB sweet-spot configs.

Original post →

More from Infra

Infra channel →