Running Qwen 27B Q4_K_M on dual RTX 3060 with llama.cpp hits ~44-50 tok/s
jacek2023 · reddit · 2026-09-26
A Redditor tested Qwen3.8-27B UD-Q4KM on a dual RTX 3060 machine using llama.cpp's llama-server with tensor parallelism, FlashAttention, 50K context, and ngram + draft-MTP speculative decoding (max 3 draft tokens). Real-world speeds: 43-50 tok/s generation and up to 198 tok/s prompt processing, making it a usable daily driver when the author's 4x3090 rig was busy with vLLM. Full command and timing logs included — evidence older 3060s still work for local LLMs.
More from Infra
- NVIDIA at $5.4 trillion is now worth more than the entire UK or French stock market — iamfakhrealam · 2026-09-26
- DeepSeek V4.1 Flash's Engram memory layer trades FFN compute for lookup tables, SemiAnalysis data suggests — teortaxesTex · 2026-09-26
- A Curated Paper List for Learning Distributed LLM Training and Inference — East-Muffin-6472 · 2026-09-26
- vLLM adds Elastic Expert Parallelism: grow/shrink MoE GPU pools under live traffic — PyTorch · 2026-09-26
- AI data center investors now favor real infrastructure over PowerPoint promises — TansuYegen · 2026-09-26
- SemiAnalysis probes DeepSeek V4.1 Flash's Engram gates, revealing how the model activates text patterns — teortaxesTex · 2026-09-26