RTX PRO 6000 Hits 2900 tok/s Running 30B Model Locally
HankYeomans · x · 2026-08-13
A developer shared benchmark results for local inference using the NVIDIA RTX PRO 6000 GPU.
- Model & Quantization: Running the Nemotron-3.5-Lightning-30B-A3B-NVFP4 model.
- Performance: Achieved an aggregated throughput of 2900 tok/s and secured 8.1M of KV cache.
- Verdict: The author notes that while the GPU is expensive, it is worth the price.
Related event: RTX PRO 6000 Hits 2900 tok/s in 30B Model Local Inference Test(2 posts)→
More from Infra
- Modular Data Center Construction to Jump from 20% to 60% by 2030, Says Bernstein — BenBajarin · 2026-08-13
- Nvidia Exec Bullish on CPO Rollout, Refuting 2029 Delay Rumors — firstadopter · 2026-08-13
- LLM Inference Explained: Why the First Token Lags and the Rest Stream Smoothly — Roger_M_Taylor · 2026-08-13
- Sony, TSMC Deal Brings Japan Chipmaking Investment to $37bn — pstAsiatech · 2026-08-13
- YMTC Overtakes Kioxia in Flash Shipments Amid AI Boom — pstAsiatech · 2026-08-13
- Robots Now Autonomously Swapping Failed Drives in Data Centers — chris_j_paxton · 2026-08-13