llama.cpp only hits 8-10 t/s on a 4bit 27B while bitsandbytes + transformers manages 27 t/s

cephaloform · x · 2026-10-10

The author reports a counterintuitive benchmark result: llama.cpp only achieves 8-10 tokens/s on a 4bit-quantized 27B model, while bitsandbytes + transformers reaches 27 tokens/s — nearly 3x faster. They are puzzled about why the popular lightweight inference stack underperforms here.

Related event: Counterintuitive Benchmark: bitsandbytes Beats llama.cpp Nearly 3x for 4-bit 27B Models(2 posts)→

Original post →

More from Infra

Infra channel →