Enabling NCCL Degrades llama.cpp Performance by 8% in Real Tests

jirka642 · reddit · 2026-08-29

A developer's real-world test on a dual RTX 3090 setup reveals that compiling llama.cpp with -DGGMLCUDANCCL=ON degrades performance (PP and TG) by approximately 8%, contrary to expectations. Despite the log suggesting NCCL would improve multi-GPU performance, disabling it yielded better speeds. The post includes full config and benchmark data.

Original post →

More from Infra

Infra channel →