Kimi K3 inference on 8× AMD MI355X hits 216 tok/s per user with TokenSpeed

zhyncs42 · x · 2026-10-10

A team reports state-of-the-art Kimi K3 inference on 8× AMD MI355X GPUs using TokenSpeed: 216 output tok/s per user at C1, 1.4× ATOM's performance, at higher precision. The deep dive covers optimizations across KDA, MLA, AttnRes, LatentMoE, and multi-GPU communication, built on portable Triton with agent-assisted Gluon specialization.

Related event: TokenSpeed Hits New SOTA Running Kimi K3 on AMD MI355X(2 posts)→

Original post →

More from Infra

Infra channel →