DeepSeek-V4-Flash Runs with 256k Context on 4x 4090 GPUs

dangerous_inference · reddit · 2026-08-05

A developer successfully deployed and ran the DeepSeek-V4-Flash model on a server with four 48GB 4090 GPUs (SM89) using DSpark and vLLM. The setup now supports a 256k context window, demonstrating a new breakthrough in running massive models on consumer-grade hardware.

Original post →

More from Infra

Infra channel →