Running DeepSeek V4 on a Single Unoptimized RTX 3090: 4 Tokens/sec

Altruistic_Heat_9531 · reddit · 2026-08-01

A developer successfully ran the DeepSeek V4 model on a single, unoptimized RTX 3090 system. Using an older CPU, an HDD, and Unsloth UD Q2 KXL quantization, they achieved a steady state of 4.02 tokens/s generation and 40 tokens/s prompt processing, demonstrating real-world local deployment metrics under heavily constrained hardware.

Original post →

More from Infra

Infra channel →