Qwen 27B on a 4060 Ti at 6-7 t/s: What Else Can Squeeze Local Inference Speed?

thatoneshadowclone · reddit · 2026-09-03

Running Qwen 27B (IQ3S Unsloth) as a coding agent on a 4060 Ti 16GB with llama.cpp, the author gets 6-7 t/s on reasoning and 11 t/s on code generation (1-1.5h per task). Current setup: dual GGUF with draft-dflash speculative decoding (max 3 drafts), -nkvo, 128K context, -fa on, q80 KV cache, b2048/ub512, full GPU offload. Asking what else can be tuned given the priority on large context.

Original post →

More from Infra

Infra channel →