GLM 5.2 Hits 2626 tok/s on a Single AMD MI355X Node

burny_tech · x · 2026-07-05

According to waferai (which hit #2 on Hacker News), engineers successfully deployed and served GLM 5.2 on AMD MI355X. The single-node throughput reached 2626 tok/s, with about 213 tok/s per user, demonstrating the potential of non-Nvidia hardware for open-source large model inference.

Related event: GLM 5.2 Achieves 2626 tok/s Inference on AMD MI355X(2 posts)→

Original post →

More from Infra

Infra channel →