RTX PRO 6000 Hits 2900 tok/s in 30B Model Local Inference Test
Developers tested the NVIDIA RTX PRO 6000 GPU for local inference, achieving an impressive 2900 tokens per second while running the Nemotron-3.5-Lightning-30B-A3B-NVFP4 model.
2026-08-13 ~ 2026-08-13 · 2 related posts
- RTX PRO 6000 Hits 2900 tok/s Running 30B Models — HankYeomans · 2026-08-13
- RTX PRO 6000 Hits 2900 tok/s Running 30B Model Locally — HankYeomans · 2026-08-13