Running Deepseek v4 Flash Q2 on a Single RTX 4090

jack_smirkingrevenge · reddit · 2026-08-15

A user reports successfully running a quantized Deepseek-v4-flash-0731 (Q2/Q3) on a single RTX 4090 with 64GB RAM. By keeping heavily utilized experts in CPU RAM and using Blaze kernels, the setup achieves a usable token rate of around 8 tps, dropping to 5 tps on disk reads due to expert misses. The author plans to publish details on this stack soon.

Original post →

More from Infra

Infra channel →