Single GH200 Runs DeepSeek V4 Flash: 2.7x Speedup with DSpark Offloading Guide

funding__secured · reddit · 2026-08-17

The author shared a complete guide for deploying DeepSeek-V4-Flash-0731 on a single NVIDIA GH200. Since the model (284B MoE) exceeds the 144GB HBM capacity, the author used vLLM's UVA (Unified Virtual Addressing) feature to offload 88GB of expert weights to the 480GB LPDDR5x system memory, combined with DSpark speculative decoding.

Key Configuration & Results:

Gotchas:

Original post →

More from Infra

Infra channel →