Surprising benchmark: bitsandbytes+transformers hits 27 t/s on 4-bit 27B vs llama.cpp's 8-10 t/s

cephaloform · x · 2026-10-10

Developer cephaloform, while trying RL for kernel optimization, found a counterintuitive result: vanilla transformers + bitsandbytes seems like the best sampling option on their hardware.

For a 4-bit quantized 27B model, bitsandbytes + transformers reaches 27 tokens/s while llama.cpp only gets 8-10 tokens/s—nearly 3x slower. The author admits this demolishes their intuition (llama.cpp is usually assumed faster) and is puzzled why. A useful real-world data point for anyone sampling on similar hardware.

Related event: Counterintuitive Benchmark: bitsandbytes Beats llama.cpp Nearly 3x for 4-bit 27B Models(2 posts)→

Original post →

More from Infra

Infra channel →