Crusoe runs 512 AMD MI355X GPUs at 5.75M tok/s in largest MLPerf inference entry

wkmyrhang · x · 2026-09-17

Crusoe's MLPerf Inference v6.1 submission served gpt-oss-120b and DeepSeek-R1 on 512 AMD Instinct MI355X GPUs — the largest MI355X entry in MLPerf history — hitting 5.75M tokens/s. Throughput scaled near-linearly from 8 to 512 GPUs at >90% of ideal over standard 400Gb Ethernet, with no InfiniBand/RoCE or cross-node RDMA needed, running as a plain Kubernetes job. Manifests are open-sourced for reproduction.

Original post →

More from Infra

Infra channel →