NVIDIA DGX Station with GB300 demos thousands of tokens per second fully local
TheZachMueller · x · 2026-09-18
NVIDIA's RTX Spark account shared a live demo by Ahmad Osman at MTS live showing the Dell Pro Max version of the NVIDIA DGX Station (GB300) running models fully locally at thousands of tokens per second, highlighting desktop-scale inference throughput for on-prem LLM workloads.
More from Infra
- Poll: What LLM gateway do you run at work when every dev holds their own API keys? — almost1it · 2026-09-18
- Dual 7900 XTX Hits 82 tok/s With RDNA3-Optimized llama.cpp Fork — deathcom65 · 2026-09-18
- OpenAI exec: GPUs hit 7-40 IQ points per watt vs human's 5, a milestone we 'zoomed past' — GregCook2011 · 2026-09-18
- China refines 91% and makes 94% of sintered rare-earth magnets powering motors from EVs to GPU data centers — demian_ai · 2026-09-18
- Mystery trader drops $100M in premium on 2-week AI stock calls expiring Oct 2 — toptickcrypto · 2026-09-18
- First-ever PyTorch Day Japan lands in Tokyo on December 10, CFP open till Sept 27 — PyTorch · 2026-09-18