DeepSeek-V4-Flash Runs with 256k Context on 4x 4090 GPUs
dangerous_inference · reddit · 2026-08-05
A developer successfully deployed and ran the DeepSeek-V4-Flash model on a server with four 48GB 4090 GPUs (SM89) using DSpark and vLLM. The setup now supports a 256k context window, demonstrating a new breakthrough in running massive models on consumer-grade hardware.
More from Infra
- Musk: Memory is the AI Bottleneck; SpaceX to Triple Nvidia Compute by 2027 — firstadopter · 2026-08-05
- Update CUDA to 13.3 to Fix DeepSeek V4 Flash Looping Issue — Easy_Werewolf7903 · 2026-08-05
- Hands-On: Deploying 550B Nemotron 3 Ultra Locally on NVIDIA DGX Station — NVIDIA Developer · 2026-08-05
- Together AI's Monthly Token Volume Skyrockets from 30B to 400T — togethercompute · 2026-08-05
- API Key Expiry Leads to Runaway Agent, Costs $300 in Idle Compute — voooooogel · 2026-08-05
- Dev Builds Pixel-Art GPU Cluster Dashboard in 20 Mins Using GLM Agent — Porespellar · 2026-08-05