Zyphra Speeds MoE Expert Routing Communication 2.63x on AMD MI300X GPUs

QuentinAnthon15 · x · 2026-10-09

Zyphra research notes that token-to-expert communication can dominate MoE model runtime. By exploiting patterns in how tokens are routed to experts, they speed up this communication by up to 2.63x on AMD MI300X GPUs — with the model itself unchanged.

Original post →

More from Infra

Infra channel →