Running DeepSeek-V4 1M Context on a Single RTX 5090 with vLLM
BlackBeardAI · reddit · 2026-08-04
A developer shared a detailed setup for running DeepSeek-V4-Flash with a full 1M context on a single RTX 5090 (32GB) paired with 256GB DDR5 RAM.
Deployment Details:
- Utilized vLLM 2.3.9 with CPU/RAM offloading, keeping most MoE experts in system RAM while retaining two complete routed MoE layers on the GPU.
- Achieved 800 tps prefill and 15+ tps decode speeds.
- Patched a底层 bug where FlashInfer's CUDA IPC hijacked TileLang.
Speculative Decoding:
- The author noted that DSpark speculative decoding behaves very differently during long reasoning sections. Draft acceptance can collapse to 30-50%, dropping generation speeds to 11-13 tok/s.
More from Infra
- NVIDIA: As AI Compute Surges, Storage and Memory Architectures Must Evolve — nordicinst · 2026-08-04
- Cloudflare Launches Programmable Wallets for AI Agents to Pay for APIs — op7418 · 2026-08-04
- Meta Introduces Helion: A New High-Level DSL for Kernel Development — ariG23498 · 2026-08-04
- DeepSeek V4 Flash 2-bit Quant Achieves 100% on Local SQL Benchmark — grumd · 2026-08-04
- 7-Month-Old Volta Raises $3B at $2.4B, Lands $10B Anthropic Deal — matt_slotnick · 2026-08-04
- Running Codex on 128-Core CPU Clusters: A New Compute Approach — BenBajarin · 2026-08-04