One RTX Pro 6K Sidecar Lifts DGX Station GB300 to 60K tok/s Prefill with DeepSeek V4.1 Flash
Sentdex · x · 2026-10-07
Sentdex shares an advanced DGX Station GB300 setup: adding an RTX PRO 6000 Blackwell Max-Q as an MoE "expert sidecar" — when a model's experts exceed the GB300's HBM, the RTX PRO 6000 holds and computes them in its own 96GB. Running DeepSeek V4.1 Flash, prefill hit 60,000 tok/s with intoxicatingly low TTFT. The referenced repo (original-el8/dgx-station-gb300-research) documents three days of measured serving results, including a MegaMoE config with hot experts on the GB300 and 3,960 cold experts on the RTX PRO 6000 (reproducing Al-ENGR's v20 recipe at 583/821 tok/s), spanning DeepSeek, MiMo, and Qwen MoE models, with every number linked to raw result files and failed quality gates flagged.
Related event: DGX Station Paired with RTX Pro 6K Hits 60K tok/s on DeepSeek Prefill(2 posts)→
More from Infra
- Macrocosmos pitches iota SDK for renting disaggregated compute across training and inference — markjeffrey · 2026-10-07
- AI neocloud Lambda raising up to $4B at $14.5B pre-money ahead of IPO — gharik · 2026-10-07
- AgentID launches: OIDC sign-in and email identity for AI agents — testingcatalog · 2026-10-07
- Agentic data toll: enterprise agent data services to hit ~$30B by 2030, says Bajarin — BenBajarin · 2026-10-07
- Huawei reportedly testing 256K-card Atlas-950 SuperPoD aiming for million-card compute — teortaxesTex · 2026-10-07
- MIT CSAIL Launches Ascent Lab, a Browser Platform for Robot Learning, With Solana Token — MIT_CSAIL · 2026-10-07