GLM 5.2 Achieves 2626 tok/s Inference on AMD MI355X

Engineers successfully deployed the open-source GLM 5.2 model on AMD MI355X accelerators, achieving a single-node throughput of 2626 tok/s and single-stream throughput of 213 tok/s. This demonstrates strong inference capabilities for open-source models on non-NVIDIA hardware.

2026-07-05 ~ 2026-07-05 · 2 related posts