vLLM on 4x vs 8x RTX 3060: does the PCIe bottleneck kill throughput?
snakeat3rr · reddit · 2026-09-25
A builder assembling a local AI server (EPYC board, Huananzhi H12D-8D, 8x16GB 2666MHz RAM) asks whether scaling from 4 RTX 3060 12GB cards at PCIe4 x16 to 8 cards at x8 (via slot bifurcation) will hurt token generation due to PCIe bandwidth limits, and whether lower-throughput 3060s suffer less than 3090s. Current baseline: 25 tps running Qwen 27B Q6 in llama.cpp layered mode on three 3060s, hoping for 50 tps with four GPUs in vLLM.
More from Infra
- LFM2.5-2.6B matches Qwen3.5-9B at 3x size, unlocking on-device agents — helloiamleonie · 2026-09-25
- Australia to host large share of next-gen datacenters, built for Anthropic — mattbeane · 2026-09-25
- Goldman: hyperscaler capex to grow 54% next year to $1.2 trillion — firstadopter · 2026-09-25
- Fine-tuning a 194M GLiNER2 dataset tagger on HF Jobs costs $1.50, lifting accuracy 10% to 69% — iamrobotbear · 2026-09-25
- A GB300 rack costs $5M at 1,580 kg — $3,165/kg, fentanyl-grade value density — StewartalsopIII · 2026-09-25
- Why GPUs need philox, not xorshift: parallel RNG in AI training explained — abhi9u · 2026-09-25