Running DeepSeek V4 on a Single Unoptimized RTX 3090: 4 Tokens/sec
Altruistic_Heat_9531 · reddit · 2026-08-01
A developer successfully ran the DeepSeek V4 model on a single, unoptimized RTX 3090 system. Using an older CPU, an HDD, and Unsloth UD Q2 KXL quantization, they achieved a steady state of 4.02 tokens/s generation and 40 tokens/s prompt processing, demonstrating real-world local deployment metrics under heavily constrained hardware.
More from Infra
- Multi-GPU Full VRAM Deployment of DeepSeek V4 Yields Only 600 t/s PP — fragment_me · 2026-08-02
- Report: SpaceX Plans 4GW Compute Cluster, Rivaling US Grid Expansion — BenBajarin · 2026-08-02
- China's AI Compute Pairings: DeepSeek x Huawei, Moonshot x Alibaba — bronzeagepapi · 2026-08-02
- Chamath Shares AI Investing Guide: Land and Power Offer Fastest Returns — davidyin44 · 2026-08-02
- Hyperscalers Report Accelerating Cloud Revenue: Google Cloud Up 82% — ai · 2026-08-02
- YC-backed Stoa launches GPU RFQ marketplace, sees $300M+ demand in first month — ycombinator · 2026-08-02