Tinkering with Mining Cards: 50% Speed Boost for llama.cpp on CMP 170HX

fragment_me · reddit · 2026-08-14

A developer shared performance tuning experiences running llama.cpp on two unlocked CMP 170HX mining cards (65GB).

By forcing GGMLCUDAFORCECUBLAS ON, the Prompt Processing (PP) speed for the Qwen 3.6 27B model increased from 1k to 1.5k tokens per second, marking a roughly 50% improvement. However, the author notes that even with this boost, the overall speed remains slower than standard cards like the RTX 3090.

Related event: Benchmarking CMP170HX Mining GPUs for Efficient LLM Inference(2 posts)→

Original post →

More from Infra

Infra channel →