AMD MI355X Cluster Hits 1M+ tokens/s on Llama 2 70B in Multi-Node Test
AntDX316 · x · 2026-08-05
AMD's MI355X GPUs demonstrate strong multi-node inference capabilities. According to test data shared by Grok, a cluster of 11 nodes (87 MI355X GPUs) running the Llama 2 70B model achieved 1,016,380 tokens/sec for online inference and 1,042,110 tokens/sec offline, with a scaling efficiency of 93–98%. Additionally, single-node performance sits at the 100k tokens/sec mark.
More from Infra
- Clever ComfyUI Node Fix Prevents MiniMax H3 OOM on 16GB VRAM — Available-Confusion2 · 2026-08-05
- a16z: Electricity Becomes AI Bottleneck as China Doubles US Generation — sujingshen · 2026-08-05
- Running LLMs on Potato PCs: Conflicting ComfyUI Flags from Different AIs — Hi7u7 · 2026-08-05
- Run MiniMax-H3 Locally: 4-bit Quantization Needs Only 8GB VRAM — petewoodbridge · 2026-08-05
- Niantic Spatial Builds City-Scale Digital Twin for Physical AI in California — petewoodbridge · 2026-08-05
- Ubuntu 26.04 LTS Introduces New HWE Stack for Confidential Computing — jedisct1 · 2026-08-05