AUR llama.cpp-cuda removal caused 10x slowdown; manual rebuild restored 1800 t/s prefill
MrHall · reddit · 2026-09-08
- A Reddit user found that after the AUR llama.cpp-cuda package was deleted, switching to the recommended ggml-cuda + llama.cpp combo (v3 optimized CachyOS builds) dropped performance to about 10% of previous speed.
- Symptom: the model loads fully onto the GPU, but processing appears to happen on the CPU.
- By recovering the old PKGBUILD and using its build script, with latest CUDA optimizations (for an RTX 5070) and CPU optimizations, prefill throughput recovered from 150 t/s to 1800 t/s.
- The poster asks why official binaries lag so far behind a custom build and complains that llama.cpp updates too often to maintain manual builds daily.
More from Infra
- FT: CXMT and YMTC stockpiled enough ASML DUV tools for three years of expansion — basedjensen · 2026-09-08
- Nvidia's $59.7B quarterly profit tops 12 iconic companies combined ($57.9B) — FinanceYF5 · 2026-09-08
- Consolidating four small models into one inference server: a doc-QA agent's ops tradeoffs — Sad-Razzmatazz-7657 · 2026-09-08
- Huawei Ascend takes PyTorch China stage: from hardware adaptation to joint standards research — PyTorch · 2026-09-08
- CPO laser series part 3: MOPA approach vs Lumentum's single-cavity route — vikramskr · 2026-09-08
- Longsys lists in Hong Kong, raising $910M as AI-driven memory boom fuels record profits — 创业邦 · 2026-09-08