Huawei’s UB-Mesh cuts switches and optics 70–80% in AI training networks
bookwormengr · x · 2026-07-22
A detailed thread on UB-Mesh, Huawei’s datacenter network architecture for AI training workloads.
- Topology options: from flat/Clos to full-mesh to hybrid designs. The current UB-Mesh-Pod reportedly uses a 4D full-mesh + Clos setup.
- Trade-off: the design reduces switch count and optical connectivity by about 70–80%, at the cost of somewhat higher latency.
- Key requirement: each NPU must be able to participate in routing, not just compute.
- The author frames it as a strong example of how Huawei is thinking about large-scale training infrastructure.
More from Infra
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11