Dual RTX 5060 Ti runs Qwen 27B: tensor parallelism lifts decode to 46 tok/s

SellToOpen · reddit · 2026-08-28

A user shares real numbers running Qwen3 27B (UD 3.0 Q4KXL) locally on two 16GB RTX 5060 Ti cards (AM4, x16+x4 lanes, 32GB DDR4) with LM Bionic. Tensor parallelism raises average decode from 28 tok/s to 46 tok/s, but drops prefill from 1000+ tok/s to 700 tok/s.

The 20k-token prompt includes writing an 800-word explanation of combustion engines, generating a Flappy Bird HTML game, and summarizing the Tolkien Wikipedia article, at default xhigh reasoning; speeds ran 22, 80-90, and 24 tok/s respectively.

MTP speculative decoding (max draft tokens 6, probability 0.88) achieved a 93.76% draft acceptance rate (9776/10427) with mean length 5.10. The author is new to local LLMs and invites corrections.

Original post →

More from Infra

Infra channel →