Signal65: one setting change on NVIDIA GB300 NVL72 adds up to 64% output throughput, 39% lower token cost
ryanshrout · x · 2026-10-02
Signal65 published its first rack-scale NVIDIA GB300 NVL72 tuning results from PINNACLE, its agentic AI benchmark: with model, nodes, and load held fixed, software-only changes deliver large gains.
- FP8 KV cache added 59–64% output throughput on large MoE models;
- A different GPU split added up to 61% on GLM-5.3-Flash — up to 39% lower cost per output token on hardware already owned.
PINNACLE measures correct work: deterministic code-scored grading against a runtime-generated answer key (no model judges), procedurally regenerated sandboxes to prevent memorization, and live multi-step agentic sessions rather than replayed traces. NVIDIA's Ian Buck and AMD both endorsed the benchmark. Takeaway: an untuned serving stack leaves enterprise rack capacity on the table.
More from Infra
- Edge0 open-sources app layer: 35B model on-device with just 1–2.5GB memory — FinanceYF5 · 2026-10-02
- SpaceX Transporter-18 launches 130 payloads including Google's orbital AI chips — XFreeze · 2026-10-02
- Top 10% of firms capture 99.5% of model-serving spend, AI compute data shows — bendee983 · 2026-10-02
- Extropic founder teases 'Thermo RSI is coming' in cryptic thermodynamic computing hype post — beffjezos · 2026-10-02
- Microsoft Foundry puts GPT-6, Claude Opus 5.5 and Grok-4.6 all in one catalog with 1M contexts — mustafasuleyman · 2026-10-02
- Musk: "Orbital compute is gonna be a very big deal" as SpaceX eyes space-based energy for AI — DimaZeniuk · 2026-10-02