Deploying 304B Model on Dual DGX Sparks: Extreme Memory Optimization

StartupTim · reddit · 2026-08-08

A developer detailed their experience deploying DeepSeek-V4-Flash-0731 (304B MoE) on 2x NVIDIA DGX Sparks (128GB unified memory each), seeking advice to free up OS RAM.

Current Setup & Bottlenecks

Ruled-Out Solutions

The Question

The author is looking for actionable ways to reduce vLLM's host-process footprint (API server + workers RSS) and tune NCCL buffers specifically for ConnectX-7 on unified memory systems.

Original post →

More from Infra

Infra channel →