Running Qwen 27B Q4_K_M on dual RTX 3060 with llama.cpp hits ~44-50 tok/s

jacek2023 · reddit · 2026-09-26

A Redditor tested Qwen3.8-27B UD-Q4KM on a dual RTX 3060 machine using llama.cpp's llama-server with tensor parallelism, FlashAttention, 50K context, and ngram + draft-MTP speculative decoding (max 3 draft tokens). Real-world speeds: 43-50 tok/s generation and up to 198 tok/s prompt processing, making it a usable daily driver when the author's 4x3090 rig was busy with vLLM. Full command and timing logs included — evidence older 3060s still work for local LLMs.

Original post →

More from Infra

Infra channel →