DDR5 Overclocking Boosts Out-of-VRAM MoE Inference by Up to 14% in Real Tests
EvolvingDior · reddit · 2026-10-06
- On an AMD 7950X with 128GB DDR5 (4x dual-rank), running MoE models that exceed VRAM via llama.cpp's customized SYCL backend on an Intel B70 32GB, the author overclocked memory from 3600MHz to 4800MHz — a measured 36% bandwidth gain.
- Results: prefill (PP) up 10%, token generation (TG) up 5%, with larger gains at deeper context: +14% pp2048 and +8.2% tg128 at 8192 context.
- Full llama-benchy tables (including prefix caching) and the complete llama-server command are provided (Qwen3.8-Flash-Next quantized GGUF + MTP speculative decoding, q80 KV cache, unified KV, 262k context), directly reproducible.
More from Infra
- Nvidia has a moat but not control: Google's TPUs now hold a quarter of global AI compute — AryHHAry · 2026-10-06
- Local LLM reality check: 4090 caps Qwen context at ~32k before VRAM runs out — BLUECOW009 · 2026-10-06
- vLLM's Transformers backend now serves text, image, audio and video with zero custom code — LysandreJik · 2026-10-06
- DeepSeek raises at least $12B, eyes IPO in early 2027, Bloomberg reports — teortaxesTex · 2026-10-06
- SEC filing: 80% of Anthropic's $518B compute commitments owed even if idle — rohanpaul_ai · 2026-10-06
- llama.cpp Metal PR boosts speculative decoding up to 3.6x on Apple Silicon — pmttyji · 2026-10-06