Running a 27B Q4 Model at 131K Context on a Single RTX 3090: ~25.6 tok/s Measured

bjivanovich · reddit · 2026-09-09

A detailed benchmark of Huihui Qwen3.8-27B Abliterated (Q4KS, Q4 KV cache) on a single RTX 3090 24GB with Xeon E5-2670 v3 and 48GB RAM:

Takeaway: a 24GB consumer GPU can hold 128K–192K context with Q4 KV cache at a steady 25 tok/s, at the cost of very long initial prefill.

Original post →

More from Infra

Infra channel →