ik_llama.cpp vs mainline: multi-GPU benchmark shows the fork 72% slower at prompt eval

vulcan4d · reddit · 2026-10-07

Reddit user vulcan4d shares counterintuitive local inference benchmarks: despite ikllama.cpp's reputation as the king of hybrid CPU/GPU offloading, mainline llama.cpp beats it across the board on a 4-GPU rig (3× P102-100 + RTX 3060) running a 177B Qwen model.

Results: Mainline hit 19.70 tok/s prompt eval and 10.86 tok/s generation, reusing 2,490 CUDA graphs; ikllama.cpp managed only 5.48 tok/s (-72%) and 8.43 tok/s (-22%), engaged zero CUDA graphs, and stalled 100ms+ on context checkpoints — it also lacks --poll support.

Key tuning tricks shared: keeping Pascal cards strictly under 9.7GB VRAM to avoid PCIe micro-paging, and using --poll to stop AVX-512 threads dropping into sleep states. The author suspects ikllama's advantage is limited to pure-CPU or single-GPU setups and asks for input from similar hybrid rigs.

Original post →

More from Infra

Infra channel →