Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU
Benchmarking shows the Bonsai quantized model outperforms Qwen quantized versions, reaching 50 tok/s at 128k context on a 24GB RTX 4090, nearly double previous speeds.
2026-09-18 ~ 2026-09-18 · 2 related posts
- Bonsai quant hits 50 tok/s at 128k context on a 24GB card, letting users run two sessions at once — julianharris · 2026-09-18
1 near-duplicate retellings: julianharris