Cascade GPU Topology Can Be Slower: The PCIe Hop Trap in Multi-GPU P2P
TheZachMueller · x · 2026-10-10
Zach Mueller surfaced a counterintuitive finding on multi-GPU communication: while Cascade topology should in theory bypass the CPU entirely for the second half, the extra PCIe switch hops make it much slower in practice — a root-based "common" topology is often faster. Mike added that with ACS disabled, p2p bandwidth opens up, and common yields better results on a single CPU.
More from Infra
- Texas data center power queue hits 474GW, 90% from data centers — FinanceYF5 · 2026-10-10
- DDR5 hits $7,200 for 256GB as engineers treat RAM as an appreciating AI asset — CtrlAltDwayne · 2026-10-10
- Lemire reruns 2026 WebSocket benchmarks: his old Bun-vs-Node.js result was wrong, Anthropic bought Bun and Cloudflare bought Deno — lemire · 2026-10-10
- Cornell/IBM paper: shared KV cache cuts looped-transformer memory 76-79% while improving quality — yoavartzi · 2026-10-10
- 4GB VRAM local LLM users: is there anything faster than llama.cpp? — your_real_Fathe_ · 2026-10-10
- Custom CUDA Megakernel Hits 140 tok/s on Qwen3.8-27B With a Single RTX 3090 — Adorable_Weakness_39 · 2026-10-10