AMD MI355X Cluster Hits 1M+ tokens/s on Llama 2 70B in Multi-Node Test

AntDX316 · x · 2026-08-05

AMD's MI355X GPUs demonstrate strong multi-node inference capabilities. According to test data shared by Grok, a cluster of 11 nodes (87 MI355X GPUs) running the Llama 2 70B model achieved 1,016,380 tokens/sec for online inference and 1,042,110 tokens/sec offline, with a scaling efficiency of 93–98%. Additionally, single-node performance sits at the 100k tokens/sec mark.

Original post →

More from Infra

Infra channel →