Dual RTX 5060 Ti runs Qwen 27B: tensor parallelism lifts decode to 46 tok/s
SellToOpen · reddit · 2026-08-28
A user shares real numbers running Qwen3 27B (UD 3.0 Q4KXL) locally on two 16GB RTX 5060 Ti cards (AM4, x16+x4 lanes, 32GB DDR4) with LM Bionic. Tensor parallelism raises average decode from 28 tok/s to 46 tok/s, but drops prefill from 1000+ tok/s to 700 tok/s.
The 20k-token prompt includes writing an 800-word explanation of combustion engines, generating a Flappy Bird HTML game, and summarizing the Tolkien Wikipedia article, at default xhigh reasoning; speeds ran 22, 80-90, and 24 tok/s respectively.
MTP speculative decoding (max draft tokens 6, probability 0.88) achieved a 93.76% draft acceptance rate (9776/10427) with mean length 5.10. The author is new to local LLMs and invites corrections.
More from Infra
- Local AI is about data ownership, not cost savings — StewartalsopIII · 2026-08-28
- SemiAnalysis: Anthropic and OpenAI to swallow majority of global compute — 新智元 · 2026-08-28
- RTX 3060 12GB: The unsung hero of local AI with 24GB VRAM and 30 t/s — I_Play_Zed · 2026-08-28
- Alibaba Open Sources Qwen3.8-Flash: Undercuts DeepSeek, Runs 1M Context on 4090 — 量子位 · 2026-08-28
- Using Langfuse traces to autonomously analyze and improve agent workflows — NielsRogge · 2026-08-28
- Nvidia arranged $500B in AI infra financing, guaranteeing $105B for OpenAI — VraserX · 2026-08-28