Cascade GPU Topology Can Be Slower: The PCIe Hop Trap in Multi-GPU P2P

TheZachMueller · x · 2026-10-10

Zach Mueller surfaced a counterintuitive finding on multi-GPU communication: while Cascade topology should in theory bypass the CPU entirely for the second half, the extra PCIe switch hops make it much slower in practice — a root-based "common" topology is often faster. Mike added that with ACS disabled, p2p bandwidth opens up, and common yields better results on a single CPU.

Original post →

More from Infra

Infra channel →