Benchmarking NUMA Performance for Local Qwen Inference

SatisfactionSuper981 · reddit · 2026-07-17

The author compared the inference performance of vLLM, ktransformers, and other solutions on NUMA systems, noting their continued shortcomings in GPU offloading and NUMA support.

They shared a custom inference engine written specifically for Qwen 3.5, featuring:

On their dual-socket Xeon 6226 machine, they provided multiple throughput metrics:

The post aims to spark community discussion on the real performance limits of NUMA combined with GPU offloading.

Original post →

More from Infra

Infra channel →