Single GH200 NVFP4 Inference Benchmark: vLLM Performance Baselines Shared
aaditya_ai · x · 2026-08-12
A developer shared inference performance baseline results running on a single NVIDIA GH200 480GB.
- Hardware & Framework: 1x GH200 running vLLM version 0.27.1.
- Test Parameters: Used NVFP4 weight precision, forced 256 output tokens/request, temperature set to 0, with median values taken from 3 runs.
- Resources: The tweet included a link to download the official NVFP4 weights provided by NVIDIA AI.
Related event: Single GH200 Hits 4600 tok/s in Nemotron vLLM Test(3 posts)→
More from Infra
- CoreWeav Adds Over $2.5B in New Customer Commitments in Early Q3 — firstadopter · 2026-08-12
- WeAreDevs Talk: Providing On-Demand Compute for AI Agents — steren · 2026-08-12
- Hetzner Launches Experimental Free LLM Inference API Featuring DeepSeek and More — AccBalanced · 2026-08-12
- Transformers.js Surpasses 10 Million Monthly Downloads, Rapid Growth Continues — nicodotdev · 2026-08-12
- d-Matrix Chip Claims 20x Speedup for Qwen Inference — TheKanter · 2026-08-12
- Starlink Offers Free Service in Colombia After Earthquake Until September 12 — DimaZeniuk · 2026-08-12