RTX PRO 6000 Hits 2900 tok/s in 30B Model Local Inference Test

Developers tested the NVIDIA RTX PRO 6000 GPU for local inference, achieving an impressive 2900 tokens per second while running the Nemotron-3.5-Lightning-30B-A3B-NVFP4 model.

2026-08-13 ~ 2026-08-13 · 2 related posts