Running DeepSeek V4 Flash on 8x 5090: 50 Concurrent Sessions for $2.5/hr
Hodler-mane · reddit · 2026-08-04
A developer shared a benchmark setup for running DeepSeek V4 Flash efficiently on 8x RTX 5090 GPUs.
- Hardware: 5090 offers the same memory bandwidth as RTX PRO 6000 and supports FP4 acceleration.
- Performance: Using the REAP mxfp4 image, it supports 50 concurrent sessions with up to 250k KV cache per session (avg 50k), delivering 30-40 tokens per second per session.
- Cost: This rig can be rented on Vast for just $2.50/hour. The author spent $150 over the weekend testing various throughput configurations, including Kimi K3.
More from Infra
- Tested: MiniMax H3 Runs Locally on 64GB MacBook Pro via Phosphene — cocktailpeanut · 2026-08-05
- NVIDIA Joins NSF Regional AI Hubs to Expand Computing Access Nationwide — nordicinst · 2026-08-05
- Agentic RL Bottlenecked by Inference: SkyPilot Halves Training Time — skypilot_org · 2026-08-05
- Hardware Architecture Debate: Why Vertical Power Delivery Over Vertical Optical IO? — jwt0625 · 2026-08-04
- CoreWeave Announces Fully Connected 2026: Fei-Fei Li & NVIDIA to Keynote — wandb · 2026-08-04
- Agentic AI Triggers a Storage Shock: Enterprise Data Becomes the New Bottleneck — BenBajarin · 2026-08-04