Bonsai Quantized Model Hits 50 tok/s with 128k Context on a 24GB GPU

Benchmarking shows the Bonsai quantized model outperforms Qwen quantized versions, reaching 50 tok/s at 128k context on a 24GB RTX 4090, nearly double previous speeds.

2026-09-18 ~ 2026-09-18 · 2 related posts

1 near-duplicate retellings: julianharris