Tinkering with Mining Cards: 50% Speed Boost for llama.cpp on CMP 170HX
fragment_me · reddit · 2026-08-14
A developer shared performance tuning experiences running llama.cpp on two unlocked CMP 170HX mining cards (65GB).
By forcing GGMLCUDAFORCECUBLAS ON, the Prompt Processing (PP) speed for the Qwen 3.6 27B model increased from 1k to 1.5k tokens per second, marking a roughly 50% improvement. However, the author notes that even with this boost, the overall speed remains slower than standard cards like the RTX 3090.
Related event: Benchmarking CMP170HX Mining GPUs for Efficient LLM Inference(2 posts)→
More from Infra
- NVIDIA and Meta Release Deployment and Sandboxed Agent Cookbook — NVIDIAAI · 2026-08-14
- NVIDIA Nemotron 3.5 on Single H200: 2k Lines of Code in 9 Secs — NVIDIAAI · 2026-08-14
- Mach33 Releases AI Compute Keystone Model Forecasting to 2040 — shaunmmaguire · 2026-08-14
- Hillock v0.4: Open-Source Neuro-Symbolic Memory Engine Running Under 1.2GB VRAM — Equivalent-Flan-1590 · 2026-08-14
- AMD Reportedly Seeking $5B Debt Financing to Accelerate AI Investment — Polymarket · 2026-08-14
- Developer Builds Offline Universe with MiniMax H3 Requiring Only 5GB VRAM — cocktailpeanut · 2026-08-14