Optimised DSv4-Flash on 2x GH200: Hits 10,000 tok/s PP via SGLang
Reddactor · reddit · 2026-08-04
A developer shared an extreme optimization scheme for running the DeepSeek v4 Flash model on dual GH200 GPUs. Using the SGLang framework and specific PRs, they achieved staggering throughput: over 10,000 tokens/s in Prefill (PP) and >300 tokens/s in Token Generation (TG). The write-up also details a nice trick to speed up Pipeline Parallelism for really long contexts.
More from Infra
- RTX 5090 Test: Minimax H3 Takes 50 Mins for 1.5MP Image — SourceTraining7959 · 2026-08-04
- Running DeepSeek V4 Flash on Mac: M3 Ultra Hits 43 tok/s — Professional-Bear857 · 2026-08-04
- Kimi K3 Speculative Decoding Model Hits 600K Downloads, Boosts AMD MI355X Throughput — bookwormengr · 2026-08-04
- Running MiniMax-H3 on RTX 3060: Renders 5s Video in 4 Minutes — Robert_Brown_7425 · 2026-08-04
- DeepSeek V4 Flash Deployed on a Single AMD MI300X GPU — zhoutong · 2026-08-04
- Arm China Goes All in AI: Launches Self-Developed NPU and Edge-to-Cloud Chips — 智东西 · 2026-08-04