DGX Station Expert Sidecar Setup Hits 5774 tok/s on DeepSeek V4.1 Flash
natesiggard · x · 2026-10-09
A hands-on setup pairs a DGX Station's GB300 with an RTX PRO 6000, running DeepSeek V4.1 Flash at 5774 tok/s. The trick: of 384 experts per layer, the 285 hottest live in GB300's HBM while the other 99 run on the RTX PRO 6000 as an expert sidecar — only tokens routed to those experts cross PCIe, and the GB300 continues if the 6000 stalls. Peak power was 970 W (GB300) and 257 W (6000). The sidecar design builds on original-el8's open GB300 research, runs on public b12x code with no unpublished kernels, and adds speculative decoding (DSpark).
More from Infra
- Amazon drops data center NDAs as community backlash spurs hundreds of moratoriums — TechCrunch AI · 2026-10-10
- Amazon drops data center NDAs, and AI agents want your credit card — TechCrunch AI · 2026-10-10
- SkyPilot founder: hoarded idle GPUs waste $20M+ a year, and the AI compute layer should be open — skypilot_org · 2026-10-10
- uv binary shrinks over 40% since July, saving nearly 4 PB of mirror bandwidth monthly — charliermarsh · 2026-10-10
- Firmus' $30B IPO collapses: only 46MW operational, valuation tripled in two months — kevinsxu · 2026-10-10
- Hyperscaler bonds now a notable share of net new Treasury borrowing — matt_slotnick · 2026-10-10