Optimised DSv4-Flash on 2x GH200: Hits 10,000 tok/s PP via SGLang

Reddactor · reddit · 2026-08-04

A developer shared an extreme optimization scheme for running the DeepSeek v4 Flash model on dual GH200 GPUs. Using the SGLang framework and specific PRs, they achieved staggering throughput: over 10,000 tokens/s in Prefill (PP) and >300 tokens/s in Token Generation (TG). The write-up also details a nice trick to speed up Pipeline Parallelism for really long contexts.

Original post →

More from Infra

Infra channel →