Upgrading to Epyc 8-channel for local LLM inference: worth it with 4 GPUs?

mrgreatheart · reddit · 2026-09-25

A Redditor with an Intel Ultra 7 system, 64GB DDR5, and four GPUs totaling 72GB VRAM (RTX 3090, 5070 Ti, two 5060 Ti) is weighing an upgrade to an Epyc 7443 + 256GB 8-channel DDR4 setup (153GB/s theoretical bandwidth) to get all cards on CPU-attached x16/x8 slots and overflow larger models to RAM. Current Qwen3.8-flash-next speeds: IQ4XS at 300 pp / 40 gen in llama.cpp, 3.05bpw at 1200 pp / 25 gen in exllamav3. They're asking whether CPU offloading large quants on Epyc is actually usable in practice.

Original post →

More from Infra

Infra channel →