Running Deepseek v4 Flash Q2 on a Single RTX 4090
jack_smirkingrevenge · reddit · 2026-08-15
A user reports successfully running a quantized Deepseek-v4-flash-0731 (Q2/Q3) on a single RTX 4090 with 64GB RAM. By keeping heavily utilized experts in CPU RAM and using Blaze kernels, the setup achieves a usable token rate of around 8 tps, dropping to 5 tps on disk reads due to expert misses. The author plans to publish details on this stack soon.
More from Infra
- Idea: AI-Generated Mini Kernels for Bare Metal Server Deployment — _Stocko_ · 2026-08-15
- DwarfStar Aims to Run Frontier Models with Low Energy — antirez · 2026-08-15
- Cambricon Revenue Doubles but Growth Slows; Valuation Pre-pays Future Expansion — 量子位 · 2026-08-15
- Free Qwen3.8-27B endpoint launched: 262K context, vision, and tool calls — victormustar · 2026-08-15
- AI infrastructure: Making 768 servers look like 1 with database sharding — blaizedsouza · 2026-08-15
- On building scalable control planes: AWS engineer deep dive — blaizedsouza · 2026-08-15