Nebius GB300 NVL72 rack tops MLPerf with 603k tokens/sec on DeepSeek R1
demian_ai · x · 2026-09-18
Nebius claimed five first-place results in MLPerf Inference v6.1: a full 72-GPU GB300 NVL72 rack topped DeepSeek R1 in both server and offline scenarios (603k and 690k tokens/sec) and pushed gpt-oss 120B past 1M tokens/sec. Nebius was one of only two submitters with preview results on NVIDIA's Vera Rubin NVL72. Scaling from 8 to 72 GPUs delivered near-linear 9x speedup.
More from Infra
- King Charles Meets OpenAI, Anthropic, DeepMind and Nvidia Execs on AI Safety — eyishazyer · 2026-09-18
- PlanetScale's TIN beats Postgres GIN full-text search: 212ms vs 288s at p99 — DanielLockyer · 2026-09-18
- NVIDIA shows 100x faster scikit-learn spectral clustering with cuML — NVIDIA Developer · 2026-09-18
- ChatGPT desktop app leaks context to the cloud by default — here's how to swap in Ollama — Technovangelist · 2026-09-18
- Reddit thread: what max-context KV reservations actually cost beyond concurrency — werunm · 2026-09-18
- Dev Patches vLLM to Run DiffusionGemma, Live Evals Show It Ties on Smarts but Loses on Speed to APIs — bodonoghue85 · 2026-09-18