Bonsai quant hits 50 tok/s at 128k context on a single 24GB GPU, up from 32
julianharris · x · 2026-09-18
The author benchmarked a Bonsai quantized model and found it faster and smaller than their favorite Qwen quant.
- On a 4090/24GB card, fully loaded 128k context runs at 50 tok/s, up from 32
- The speedup allows two concurrent sessions on one 24GB GPU
- 256k full context is estimated at roughly 25 tok/s
- A "drafted with Bonsai" variant could push performance further, though the author hasn't located one yet
Related event: Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU(2 posts)→
More from Infra
- Hedge fund CIO: Anthropic burns maybe 80% less per token than OpenAI — rohanpaul_ai · 2026-09-18
- Engineer: companies hide servers from management to dodge forced cloud migrations — irth7 · 2026-09-18
- SK Hynix subsidiary Solidigm plans to build a NAND flash fab in the US — zephyr_z9 · 2026-09-18
- Huawei says China's AI training will shift to Ascend 950DT SuperPoDs starting 2027 — pstAsiatech · 2026-09-18
- Jev playground shows inference and roundtrip latency; EU users pay 120ms extra — DanielLockyer · 2026-09-18
- HEIF Heist: libheif flaws allowed researchers to hack OpenAI, Slack, Meta and more — tszzl · 2026-09-18