Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs
QuentinAnthon15 · x · 2026-10-09
Zyphra research notes that token-to-expert communication can dominate MoE model runtime. By exploiting patterns in how tokens are routed to experts, they speed up this communication by up to 2.63x on AMD MI300X GPUs — with the model itself unchanged.
More from Infra
- Tencent's STEPQuant: 6-bit recurrent states match FP32 with 68.7% less memory — _akhaliq · 2026-10-09
- Bain projects 38.6M GPU and custom silicon shipments by 2030, 10x 2023's 3.9M — Beth_Kindig · 2026-10-09
- AWS reference architecture: multi-team GPU cluster sharing on SageMaker HyperPod — AWS ML Blog · 2026-10-09
- Mistral slammed for training open models on datacenters powered ~70% by coal — wavefnx · 2026-10-09
- Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate — Vasili_Sk · 2026-10-09
- NVIDIA's NeMo-DCR cuts 1T-model weight sync from 87.5 min to 150s, 12-40x faster checkpoint transfer — dair_ai · 2026-10-09