Microsoft HotChips presentation reveals bandwidth wall in AI cluster architecture

bookwormengr · x · 2026-08-31

At HotChips, a Microsoft presentation highlighted a "bandwidth wall" in AI cluster architecture: bandwidth between local accelerators is 6-7 400G (non-switched), while between different compute trays/racks it drops to 2 400G (switched). Although Microsoft uses a unified protocol and on-chip NICs, this disparity is significant. Technically, this wall could be removed by redistributing the 28 400G links per accelerator, depending on tensor parallelism requirements beyond 4 nodes.

Original post →

More from Infra

Infra channel →