Building a $15k Local Inference Rig for DeepSeek V4 Flash: MI210 vs A40?
_TheWolfOfWalmart_ · reddit · 2026-09-12
A Reddit user seeks GPU advice for a $15k work rig to run DeepSeek V4 Flash and similar models locally with at least 128GB VRAM and minimal quantization, fitting 3 cards in a Dell R740. Candidates: 3× AMD MI210 (192GB) vs 3× NVIDIA A40 (144GB); they're asking for real token-gen and prompt-processing benchmarks and say cloud rental for testing is unavailable.
More from Infra
- Instinct may burn $100M+ a year in tokens, and open-weight models aren't actually cheaper — ivan_bezdomny · 2026-09-12
- Curie: a from-scratch 17B model designed to run from SSD, 33 tokens/s on one CPU core — Just_Vugg_PolyMCP · 2026-09-12
- LMStudio now accepts llama.cpp overrides; --yarn-attn-factor 1.2 may boost creativity — Extraaltodeus · 2026-09-12
- Analyst details Apple's S11, A20 and M6 silicon: new packaging, cooling and on-device AI designs — BenBajarin · 2026-09-12
- Custom Silicon 3.0: market shifts to program responsibility as agentic AI enters cyber defense — BenBajarin · 2026-09-12
- Single 3090 + 32GB RAM: squeezing max fidelity out of local Qwen3.8-27B — Certain_Yam_5824 · 2026-09-12