Moonshot inference may require GB200-class clusters

zephyr_z9 · x · 2026-07-20

A post arguing that Moonshot must be running GB200s to serve the model efficiently at inference time. The quote says the model would not fit on a single Hopper node, and that multi-node serving on non-NVL72 systems would be inefficient.

The attached image shows a large-scale compute rack setup labeled “Atlas 950 SuperPod” with visible specs including 256TB unified memory, 1024 GPUs, and 3-split RT. The overall takeaway is that the discussion is about the compute stack required for efficient inference at very large scale.

Original post →

More from Infra

Infra channel →