For LLM workloads, AMD CCD count matters more than core count: 16-core tops out at 125GB/s
HankYeomans · x · 2026-09-24
A hands-on lesson for LLM hardware builds: core count isn't the bottleneck — AMD CCD count is. Eight slots of CL32/6400MT memory deliver 409GB/s bandwidth, but a 16-core Threadripper PRO has only two CCDs, capping the CPU critical path at 125GB/s (140GB/s with fabric overclock). To actually use the full memory bandwidth on CPU-side workloads, you need a 64-core part with 8 CCDs.
More from Infra
- A Ready-to-Use Prompt That Makes Your Agent Audit Its Own API Bills — gethackteam · 2026-09-24
- ~50us per kernel launch possible, but only by forking a custom single-model inference stack — AlpinDale · 2026-09-24
- Tencent Hunyuan extends critical-batch-size theory to LLM RL: 29% faster GRPO, 2.29× PPO throughput — TencentHunyuan · 2026-09-24
- Xiaomi Previews HySparse2 for MiMo-V3: KV Cache Cut to ~1/4.5 at 1M Context — 智东西 · 2026-09-24
- Open-source ComfyUI nodes losslessly compress models: 28GB to 19GB, bit-identical — New-Shift6661 · 2026-09-24
- NVIDIA demos Nemotron 3.5 Lightning running locally on DGX Spark and Station — NVIDIA Developer · 2026-09-24