ik_llama.cpp vs mainline: multi-GPU benchmark shows the fork 72% slower at prompt eval
vulcan4d · reddit · 2026-10-07
Reddit user vulcan4d shares counterintuitive local inference benchmarks: despite ikllama.cpp's reputation as the king of hybrid CPU/GPU offloading, mainline llama.cpp beats it across the board on a 4-GPU rig (3× P102-100 + RTX 3060) running a 177B Qwen model.
Results: Mainline hit 19.70 tok/s prompt eval and 10.86 tok/s generation, reusing 2,490 CUDA graphs; ikllama.cpp managed only 5.48 tok/s (-72%) and 8.43 tok/s (-22%), engaged zero CUDA graphs, and stalled 100ms+ on context checkpoints — it also lacks --poll support.
Key tuning tricks shared: keeping Pascal cards strictly under 9.7GB VRAM to avoid PCIe micro-paging, and using --poll to stop AVX-512 threads dropping into sleep states. The author suspects ikllama's advantage is limited to pure-CPU or single-GPU setups and asks for input from similar hybrid rigs.
More from Infra
- Paired 4:8 sparsity gives 1.35-1.65x over dense NVFP4 on B200, 1.18x in serving — vllm_project · 2026-10-07
- Same GPUs, 100x Gap: vLLM Hits 110 tok/s on Dual RTX 5090s Where llama.cpp Took 30 Minutes — vllm_project · 2026-10-07
- kipply's July-August digest: export controls lifted, alignment-faking paper, and more — kipperrii · 2026-10-07
- Marvell Investor Day 2026 materials land, with a nudge to fix the chart arrows — jwt0625 · 2026-10-07
- Cloud credits in a bubble: AWS clones, Microsoft billing traps, Google's dead services — mkheck · 2026-10-07
- NVFP4 runs a 2.4T-parameter model on a quarter of the GPUs, cutting per-token cost up to 73% — ryanshrout · 2026-10-07