DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station

antirez · x · 2026-08-17

Antirez demonstrated running the DeepSeek v4 PRO Q2 model on a DGX Station with optimized inference. By splitting routed experts between VRAM and RAM and optimizing kernels for this hybrid setup, the system achieved 45 t/s, with potential for further speedups. The author highlights this as a demonstration of DwarfStar's capabilities as a workstation inference engine.

Related event: Redis Author Optimizes DeepSeek V4 to 45 t/s(3 posts)→

Original post →

More from Infra

Infra channel →