Upgrading to Epyc 8-channel for local LLM inference: worth it with 4 GPUs?
mrgreatheart · reddit · 2026-09-25
A Redditor with an Intel Ultra 7 system, 64GB DDR5, and four GPUs totaling 72GB VRAM (RTX 3090, 5070 Ti, two 5060 Ti) is weighing an upgrade to an Epyc 7443 + 256GB 8-channel DDR4 setup (153GB/s theoretical bandwidth) to get all cards on CPU-attached x16/x8 slots and overflow larger models to RAM. Current Qwen3.8-flash-next speeds: IQ4XS at 300 pp / 40 gen in llama.cpp, 3.05bpw at 1200 pp / 25 gen in exllamav3. They're asking whether CPU offloading large quants on Epyc is actually usable in practice.
More from Infra
- Lambda engineer shares local inference build rule: 27B models need 24-32GB VRAM — TheZachMueller · 2026-09-25
- Pokee AI demos 36B agent model running fully local on Snapdragon X2 Elite with 32GB RAM — Kyrannio · 2026-09-25
- AMD to present MXFP8 pretraining scaling on 1K+ MI355X GPUs at PyTorchCon 2026 — PyTorch · 2026-09-25
- AI energy startup Parallax launches with $117m from Founders Fund, Lux, Greylock and others — graceisford · 2026-09-25
- Nebius/WEKA benchmark: shared KV cache lifts agentic inference throughput 2.4x with 93% hit rate — AccBalanced · 2026-09-25
- Burkov's TP Weekly #179: GPU rent vs buy, llm-d serving 753B model at 5-10x lower cost — burkov · 2026-09-25