llama.cpp only hits 8-10 t/s on a 4bit 27B while bitsandbytes + transformers manages 27 t/s
cephaloform · x · 2026-10-10
The author reports a counterintuitive benchmark result: llama.cpp only achieves 8-10 tokens/s on a 4bit-quantized 27B model, while bitsandbytes + transformers reaches 27 tokens/s — nearly 3x faster. They are puzzled about why the popular lightweight inference stack underperforms here.
More from Infra
- Tinker Cuts RL Training Prices Up to 70%, Adds GLM-5.3-Flash and DeepSeek-v4.1-Flash — soumithchintala · 2026-10-10
- Hyperscaler bonds now rival Treasury borrowing as AI capex shifts from cash to debt — QuintinPope5 · 2026-10-10
- Ben Bajarin: AI demand runs 115-120% above capacity, 2027 to be peak constraint year — BenBajarin · 2026-10-10
- AMD at $1T trades richer than NVIDIA on both multiples — analyst warns of a stall, 320M warrants loom — pdamodaran · 2026-10-10
- Two H100 price indices show just 0.17 weekly correlation, clouding compute futures hedge — BenBajarin · 2026-10-10
- Chrome's new echo canceller halves voice agent word error rate, stops agents answering their own greeting — chadwallacehart · 2026-10-10