Bonsai quant hits 50 tok/s at 128k context on a single 24GB GPU, up from 32

julianharris · x · 2026-09-18

The author benchmarked a Bonsai quantized model and found it faster and smaller than their favorite Qwen quant.

Related event: Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU(2 posts)→

Original post →

More from Infra

Infra channel →