Surprising benchmark: bitsandbytes+transformers hits 27 t/s on 4-bit 27B vs llama.cpp's 8-10 t/s
cephaloform · x · 2026-10-10
Developer cephaloform, while trying RL for kernel optimization, found a counterintuitive result: vanilla transformers + bitsandbytes seems like the best sampling option on their hardware.
For a 4-bit quantized 27B model, bitsandbytes + transformers reaches 27 tokens/s while llama.cpp only gets 8-10 tokens/s—nearly 3x slower. The author admits this demolishes their intuition (llama.cpp is usually assumed faster) and is puzzled why. A useful real-world data point for anyone sampling on similar hardware.
More from Infra
- Industrially, What You Do More Of Gets Cheaper: The Learning Curve Behind Data Center Costs — Afinetheorem · 2026-10-10
- SGLang on NVIDIA Vera Rubin: up to 20% faster Kimi K3 inference, 4.8x gains for Cognition — ying11231 · 2026-10-10
- Agent-built system hits 2,242 tok/s on AMD MI300As, 2.33× faster than SGLang in 105 hours — bariskasikci · 2026-10-10
- Bespoke serving systems becoming the only viable path as hardware-model combos explode — bariskasikci · 2026-10-10
- Napkin math on the AI bubble: 100x cheaper models still mean 1000x more GPUs — gabriel1 · 2026-10-10
- Super Micro contractor pleads guilty in $2.5B scheme smuggling Nvidia AI chips into China — Polymarket · 2026-10-10