Qwen3.8-27B hits >70 tok/s and full 262k context on 2x3090 with vanilla vLLM

maqifrnswa · reddit · 2026-09-22

On 2x3090s, a user benchmarks Qwen3.8-27B-INT4 (RedHatAI quant) with vanilla vLLM: >70 tok/s single-stream decode, 200 tok/s at 3-4 concurrency, >10k tok/s prefill, and full 262k context. Key settings: INT4 quant ideal for Ampere, weighted fp8 KV cache (near-identical to bf16), MTP speculative decoding (3 tokens), tensor parallel 2, prefix caching. Tuned over a month using vLLM metrics + Prometheus and vllm bench serve; full serve command posted. The author argues the overlooked RedHat quant is ideal for Ampere.

Original post →

More from Infra

Infra channel →