Running DeepSeek at 1M Context on Single RTX 5090 via Adaptive Dual CUDA Graphs

BlackBeardAI · reddit · 2026-08-05

A developer shared the latest optimizations for running the massive DeepSeek-V4-Flash-0731 (155GB) on a single RTX 5090 while maintaining its native 1M token context.

Previously, fixed speculative decoding depth caused unstable throughput due to varying acceptance rates between hard reasoning and code generation. To fix this, the author developed a phase-adaptive mechanism:

To prevent performance drops from mismatched CUDA graph shapes during switching, the author implemented dual CUDA graph capture. Ultimately, at the full 1M context, reasoning throughput reached 13.8 tok/s and code generation hit 17.0 tok/s.

Related event: Single RTX 5090 Runs DeepSeek 1M Context(2 posts)→

Original post →

More from Infra

Infra channel →