llama.cpp NUMA mirroring boosts dual-EPYC inference by up to 137%
mattescala · reddit · 2026-08-30
Addressing the dual-socket bottleneck where cores read weights across the slow interconnect, the author enabled the unused NUMAMIRROR strategy in ggml-cpu.h. This keeps a full copy of weights on each NUMA node, doubling RAM usage but allowing local access. Benchmarks show 64-71% speedups for DeepSeek-V4 and GLM-5.2, and a massive 137% boost for gemma-4-31B dense decoding. A PR has been submitted to llama.cpp.
More from Infra
- TensorSharp integrates MiniMax H3 for local image-to-video inference — fuzhongkai · 2026-08-30
- Qwen3.8-Flash-Next on 2x DGX Spark NVFP4: 50 t/s decode, 2,900 t/s prefill — -dysangel- · 2026-08-30
- SEC filings this week: 50MW Microsoft compute deal, banks roll out agentic AI workspaces — Justgototheeffinmoon · 2026-08-30
- Quantized Qwen 3.8 Flash Next may be best local model for 64GB Macs — antirez · 2026-08-30
- AutoNodo runs 28 days processing 14B tokens, exploring massive context engineering — nodo48 · 2026-08-30
- SemiAnalysis: OpenAI Wins on Both Speed and Cost, Beating Nvidia and Startups — firstadopter · 2026-08-30