Intel Xeon LLM Bandwidth Halved? Local MoE Inference Tuning Log
GetOutOfMyFeedNow · reddit · 2026-08-09
A developer attempted to run DeepSeek-V4-Flash-0731 (MXFP4 format) on a Xeon w7-3465, expecting a theoretical max memory bandwidth of 153GB/s, but actual performance fell significantly short.
- Hardware: 4x RDIMM DDR5-4800, 28 physical cores, 56 logical cores.
- Bottleneck: Measured bandwidth hovered around 36-40GB/s, with inference crawling at 3-4 tokens/s.
- Tuning: Assisted by GPT, tweaking --threads and --threads-batch to match physical cores yielded a 2.4x speedup in one iteration, but overall performance still hasn't reached the machine's theoretical limits.
More from Infra
- Replace ChatGPT Plus with Local Models: A 5-Step Guide — Aiden_Tech_Ai · 2026-08-09
- Nvidia's Rubin Ultra Shifts from HBM to Optical Interconnects, Altering Market Dynamics — zephyr_z9 · 2026-08-09
- Running SD Natively on Android: SDXL Takes 20 Minutes on a Phone — Silent-Paramedic4063 · 2026-08-09
- LFM 2.6B Hits 260 Tokens/s on RTX 3090: A Dev's Hands-On Review — Borkato · 2026-08-09
- Report: Nvidia to Invest Up to $3B in AI Data Center Power Developer Lancium — rohanpaul_ai · 2026-08-09
- Running an LLM on an ESP32 with Only 81KB of Memory — Similar_Wealth_1850 · 2026-08-09