DeepSeek v4 PRO Q2 hits 45 t/s with VRAM/RAM split on DGX Station
antirez · x · 2026-08-17
Antirez demonstrated running the DeepSeek v4 PRO Q2 model on a DGX Station with optimized inference. By splitting routed experts between VRAM and RAM and optimizing kernels for this hybrid setup, the system achieved 45 t/s, with potential for further speedups. The author highlights this as a demonstration of DwarfStar's capabilities as a workstation inference engine.
Related event: Redis Author Optimizes DeepSeek V4 to 45 t/s(3 posts)→
More from Infra
- PufferLib prerelease supports >20M steps/second reinforcement learning — jsuarez · 2026-08-17
- Data center boom drives up emissions as costs outweigh benefits — robleclerc · 2026-08-17
- Model routing should be optimized at the harness layer, not the gateway — agihouse_org · 2026-08-17
- Qwen3.8-27B Benchmarks: RPC inference tests across AMD and Nvidia GPUs — tabletuser_blogspot · 2026-08-17
- AI Spending Gap 625x: Top 1% Spend $7,500/Employee/Month vs Median $12 — rohanpaul_ai · 2026-08-17
- AWS to deploy Nvidia GB300 as primary GPU in 2026, expand Trainium shipments: TrendForce — Beth_Kindig · 2026-08-17