Nebius GB300 NVL72 rack tops MLPerf with 603k tokens/sec on DeepSeek R1

demian_ai · x · 2026-09-18

Nebius claimed five first-place results in MLPerf Inference v6.1: a full 72-GPU GB300 NVL72 rack topped DeepSeek R1 in both server and offline scenarios (603k and 690k tokens/sec) and pushed gpt-oss 120B past 1M tokens/sec. Nebius was one of only two submitters with preview results on NVIDIA's Vera Rubin NVL72. Scaling from 8 to 72 GPUs delivered near-linear 9x speedup.

Original post →

More from Infra

Infra channel →