NVFP4 runs a 2.4T-parameter model on a quarter of the GPUs, cutting per-token cost up to 73%

ryanshrout · x · 2026-10-07

Signal65 PINNACLE rack-scale testing shows serving a 2.4-trillion-parameter model in BF16 took 16 of 18 nodes on an NVIDIA GB300 NVL72 rack; with NVFP4 the same model ran on 4 nodes at 93% of BF16 throughput.

That works out to 3.7x the work per GPU and up to 73% lower cost per output token. The author's takeaway for anyone planning to run frontier open models on their own hardware: the GPUs a model occupies determine true cost, and NVFP4 can cut that number substantially as Signal65 PINNACLE moves to rack scale with NVIDIA.

Original post →

More from Infra

Infra channel →