Running a 27B Q4 Model at 131K Context on a Single RTX 3090: ~25.6 tok/s Measured
bjivanovich · reddit · 2026-09-09
A detailed benchmark of Huihui Qwen3.8-27B Abliterated (Q4KS, Q4 KV cache) on a single RTX 3090 24GB with Xeon E5-2670 v3 and 48GB RAM:
- 131072 context: 111,797 tokens used (85.3%), VRAM at 23.5/24GB; first run 22.02 tok/s with a 303s TTFT, subsequent runs average 25.58 tok/s with TTFT between 4.7–50s
- 196608 context: average 25.80 tok/s, cold prefill up to 514s
- Average DTA 64.57%
Takeaway: a 24GB consumer GPU can hold 128K–192K context with Q4 KV cache at a steady 25 tok/s, at the cost of very long initial prefill.
More from Infra
- Cerebras CTO's chip architecture deep dives—WSE-3, Hot Chips 34, Cornell lectures—barely get any views — blaizedsouza · 2026-09-09
- Compute partners want 40% prepaid 5-year auction terms when you're struggling—then flock to 'help' once you succeed — saranormous · 2026-09-09
- OpenAI plans ~$1T compute spend; expert sees it fueling AGI, not inference — NinaDSchick · 2026-09-09
- Compute markets split: financial players want standardized exposure, operators want bespoke delivery — AccBalanced · 2026-09-09
- Humanoid revolution needs $250B more motors on top of $50-60B in use today — stepjamUK · 2026-09-09
- Google Cloud breaks down 4 ways to serve open models, from managed to fully yours — rseroter · 2026-09-09