Dual 3090 Qwen 27B Full Context Setup: 3200 tok/min Achieved

CryptographerLow7817 · reddit · 2026-08-23

Author shared a production-ready setup for running Qwen3.8-27B with 262k context on dual RTX 3090s (no NVLink). Using vLLM 0.27.1, MTP-3 speculative decoding, and GPU prefix caching, a 96.8% cache hit rate was achieved. Performance: Mean TTFT 6.5s (vs 140s cold), single-session decode at 80-90 tok/s, and multi-session throughput of 3200 tok/min.

Original post →

More from coding & agent

coding & agent channel →