vLLM on 4x vs 8x RTX 3060: does the PCIe bottleneck kill throughput?

snakeat3rr · reddit · 2026-09-25

A builder assembling a local AI server (EPYC board, Huananzhi H12D-8D, 8x16GB 2666MHz RAM) asks whether scaling from 4 RTX 3060 12GB cards at PCIe4 x16 to 8 cards at x8 (via slot bifurcation) will hurt token generation due to PCIe bandwidth limits, and whether lower-throughput 3060s suffer less than 3090s. Current baseline: 25 tps running Qwen 27B Q6 in llama.cpp layered mode on three 3060s, hoping for 50 tps with four GPUs in vLLM.

Original post →

More from Infra

Infra channel →