NVFP4 runs a 2.4T-parameter model on a quarter of the GPUs, cutting per-token cost up to 73%
ryanshrout · x · 2026-10-07
Signal65 PINNACLE rack-scale testing shows serving a 2.4-trillion-parameter model in BF16 took 16 of 18 nodes on an NVIDIA GB300 NVL72 rack; with NVFP4 the same model ran on 4 nodes at 93% of BF16 throughput.
That works out to 3.7x the work per GPU and up to 73% lower cost per output token. The author's takeaway for anyone planning to run frontier open models on their own hardware: the GPUs a model occupies determine true cost, and NVFP4 can cut that number substantially as Signal65 PINNACLE moves to rack scale with NVIDIA.
More from Infra
- Marvell investor day: questions raised over 100mm die vs future 300mm substrates — jwt0625 · 2026-10-07
- Cloud credits in a bubble: AWS clones, Microsoft billing traps, Google's dead services — mkheck · 2026-10-07
- ik_llama.cpp vs mainline: multi-GPU benchmark shows the fork 72% slower at prompt eval — vulcan4d · 2026-10-07
- Intel exec: agentic AI spends real time waiting on business systems, end-to-end wall-clock is the metric that matters — ryanshrout · 2026-10-07
- Serving & Monitoring Local LLMs Across Three Mixed-GPU Machines Without Duct Tape — ziyaulhuk12 · 2026-10-07
- Disaggregated inference is the future, says e/acc's Beff Jezos after panel with General Compute — beffjezos · 2026-10-07