Benchmarking NUMA Performance for Local Qwen Inference
SatisfactionSuper981 · reddit · 2026-07-17
The author compared the inference performance of vLLM, ktransformers, and other solutions on NUMA systems, noting their continued shortcomings in GPU offloading and NUMA support.
They shared a custom inference engine written specifically for Qwen 3.5, featuring:
- KV cache fully stored on GPU
- Dense and shared experts also placed on GPU
- Experimental support for partial offloading from the CPU
On their dual-socket Xeon 6226 machine, they provided multiple throughput metrics:
- Qwen 3.5 35B: decode at 45 t/s, Nmoe reaching 71 t/s
- Qwen 3.5 122B: decode at 20–40 t/s, depending on the engine
- Qwen 3.5 397B: runs at 15–17 t/s on CPU after encountering CUDA OOM
The post aims to spark community discussion on the real performance limits of NUMA combined with GPU offloading.
More from Infra
- oMLX 0.5.2 adds Mac menu-bar stats, low-bit decode kernels, and faster downloads — awnihannun · 2026-07-22
- Strangeworks launches Aura to turn enterprise ops into production optimization systems — whurley · 2026-07-22
- Graph workload 854.graph500 enters SPEC CPU 2026 as a new CPU benchmark — Prof_DavidBader · 2026-07-22
- Hybrid and local inference are emerging as a response to AI energy and token costs — dmitry140 · 2026-07-22
- NVIDIA details Vera CPU with 2x performance claims and a 22,000-core rack — ryanshrout · 2026-07-22
- NVIDIA says Vera Rubin NVL72 delivers 10x more tokens per megawatt than Blackwell — nvidia · 2026-07-22