Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once

julianharris · x · 2026-09-18

julianharris reports that the Bonsai quant is faster and smaller than his favorite Qwen quant: 50 tok/s at 128k fully loaded on a 4090/24GB (vs 32 before) — fast enough to run TWO sessions at once on one card. He expects roughly 25 tok/s at 256k filled context. Local long-context inference on consumer GPUs has improved significantly.

Related event: Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU(2 posts)→

Original post →

More from Infra

Infra channel →