Dual 5060Ti Inference: Is PCIe Bandwidth a Bottleneck?
Ok-Conflict391 · reddit · 2026-08-21
A user running a 27B parameter model (Q6 quantized) on dual RTX 5060Ti GPUs achieved 40-50 tps. HWinfo monitoring revealed that both the PCIe 5x16 and PCIe 4x4 slots were fully saturated during inference. The user suspects the second x4 slot is a bottleneck and is considering moving the GPU to a spare PCIe 5x4 NVMe slot via a riser.
More from Infra
- Proposed London Datacenter to Emit 1m Tonnes of CO2 Annually — nordicinst · 2026-08-21
- Running Firecracker on M-series Macs via Nested Virtualization — dejavucoder · 2026-08-21
- MLX-LM maintainer on project focus and governance — andrejusb · 2026-08-21
- GCC and Clang Leave 64-bit Division on Fast Path at -O2 — tetsuoai · 2026-08-21
- AI eats the DRAM supply: DDR5 prices up 500% in a year, back to 2007 levels — 量子位 · 2026-08-21
- Nix-like setup in TypeScript: Managing multi-machine configs with AI — samgoodwin89 · 2026-08-21