Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested
No_Run8812 · reddit · 2026-09-10
A software developer asks Reddit to review his local AI server upgrade: two RTX Pro 6000 GPUs, plus newly ordered 128GB DDR5-5600 quad-channel memory, an AMD Threadripper PRO 9955W, a 360mm AIO cooler and an ASUS Pro WS WRX90E-SAGE SE motherboard. He plans to run DeepSeek V4 Flash and Qwen 3.8 Flash, and asks whether GLM 5.3 Flash with RAM offload could achieve decent speeds on this rig, inviting the community to flag any obvious build mistakes.
More from Infra
- Kimi K3 lands on RunPod: 2.8T params, 1M context, $3/$15 per 1M tokens — Kimi_Moonshot · 2026-09-10
- Dev burns 300M tokens on GLM 5.3 in a week and still has quota left — saibharadwaj · 2026-09-10
- After Nvidia's Hugging Face buyout, devs call for a neutral alternative — hargup13 · 2026-09-10
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Screenshot surfaces rare admission of 72-hour KV cache limits in V4-era architecture — zephyr_z9 · 2026-09-10