GLM 5.2 Hits 2626 tok/s Inference on AMD MI355X

SumitGup · x · 2026-07-05

An engineer showcased serving the GLM 5.2 model on AMD MI355X accelerators, achieving a single-node throughput of 2626 tok/s and a single-stream throughput of 213 tok/s. This demonstrates the competitive inference optimization capabilities of non-NVIDIA hardware for mainstream open-source large models.

Related event: GLM 5.2 Achieves 2626 tok/s Inference on AMD MI355X(2 posts)→

Original post →

More from Infra

Infra channel →