Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once
julianharris · x · 2026-09-18
julianharris reports that the Bonsai quant is faster and smaller than his favorite Qwen quant: 50 tok/s at 128k fully loaded on a 4090/24GB (vs 32 before) — fast enough to run TWO sessions at once on one card. He expects roughly 25 tok/s at 256k filled context. Local long-context inference on consumer GPUs has improved significantly.
Related event: Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU(2 posts)→
More from Infra
- Who Opens Up the Compute Middle Market? The $10-100M Deployment Standards Layer — AccBalanced · 2026-09-18
- Announced data centers hit 342GW vs 50GW installed, implying $26.3T in spend — BenBajarin · 2026-09-18
- Hyperbolic hires quant researchers to build GPU compute as a tradable asset class — YiMaTweets · 2026-09-18
- Crusoe raises $3.9B Series F at $30.9B valuation to fuel AI energy buildout — beffjezos · 2026-09-18
- Periodic Labs details its stack: 4.1x Megatron throughput, frontier-beating science models on 1,300 H200s — hsu_byron · 2026-09-18
- Jeff Dean: a handful of workloads will dominate world compute, 'crying out' for specialized silicon — AccBalanced · 2026-09-18