Running DeepSeek-V4 1M Context on a Single RTX 5090 with vLLM

BlackBeardAI · reddit · 2026-08-04

A developer shared a detailed setup for running DeepSeek-V4-Flash with a full 1M context on a single RTX 5090 (32GB) paired with 256GB DDR5 RAM.

Deployment Details:

Speculative Decoding:

Original post →

More from Infra

Infra channel →