Enabling NCCL Degrades llama.cpp Performance by 8% in Real Tests
jirka642 · reddit · 2026-08-29
A developer's real-world test on a dual RTX 3090 setup reveals that compiling llama.cpp with -DGGMLCUDANCCL=ON degrades performance (PP and TG) by approximately 8%, contrary to expectations. Despite the log suggesting NCCL would improve multi-GPU performance, disabling it yielded better speeds. The post includes full config and benchmark data.
More from Infra
- Open Source Static Performance Model for LLM Inference — stanfordnlp · 2026-08-29
- Prediction: Closed frontier models to become downloadable by 2027 — imjustnewatai · 2026-08-29
- Achieving 181 tok/s on Qwen3.8 with 2x DGX Sparks via NVMe offloading — StartupTim · 2026-08-29
- Together AI processes 135B+ GLM-5.3 Flash tokens in 24 hours — togethercompute · 2026-08-29
- Woof introduces liquidity to AI compute via onchain securitization — edgarpavlovsky · 2026-08-29
- Nvidia strategy: Buy the open-source layer underneath — bindureddy · 2026-08-29