Counterintuitive Benchmark: bitsandbytes Beats llama.cpp Nearly 3x for 4-bit 27B Models
A developer's benchmark found that running a 4-bit 27B model with transformers + bitsandbytes reached 27 tokens/s, nearly 3x faster than llama.cpp's 8-10 tokens/s on the same hardware, challenging common assumptions about quantized inference speed.
2026-10-10 ~ 2026-10-10 · 2 related posts
- llama.cpp only hits 8-10 t/s on a 4bit 27B while bitsandbytes + transformers manages 27 t/s — cephaloform · 2026-10-10
- Surprising benchmark: bitsandbytes+transformers hits 27 t/s on 4-bit 27B vs llama.cpp's 8-10 t/s — cephaloform · 2026-10-10