256GB 8-Channel Epyc Inference Rig: Is a Single 3090 Worth Adding?
u_Leon · reddit · 2026-09-17
A Redditor is building a second-hand 256GB DDR4 8-channel Epyc rig (Asrock ROMED8-2T) to run large models cheaply, costing less than a new 128GB Strix machine. The question: would adding a single 3090 meaningfully improve prefill for 200GB models like GLM-5.3-Flash at 4-bit, versus the much pricier and power-hungry 4x 3090 setup?
More from Infra
- Tencent open-sources FlexKV distributed KV cache for LLM inference, cutting TTFT by up to 70% — Roger_M_Taylor · 2026-09-17
- NVIDIA releases NVFP4 quantized DeepSeek-V4.1-Flash on Hugging Face — TheZachMueller · 2026-09-17
- Optimization mined via Bittensor competition lands in vLLM, boosting Qwen3 throughput ~4% — const_reborn · 2026-09-17
- Early vLLM PR adds Jev-like structured generation for DiffusionGemma, only 2x endpoint latency on a DGX Spark — generativist · 2026-09-17
- New inference engine Atlas debuts, redditor says it beats llama.cpp on Strix Halo — einthecorgi2 · 2026-09-17
- Keeping vLLM's prefix cache warm between agent turns: an engineering guide — bolts98 · 2026-09-17