Qwen3.8-27B hits >70 tok/s and full 262k context on 2x3090 with vanilla vLLM
maqifrnswa · reddit · 2026-09-22
On 2x3090s, a user benchmarks Qwen3.8-27B-INT4 (RedHatAI quant) with vanilla vLLM: >70 tok/s single-stream decode, 200 tok/s at 3-4 concurrency, >10k tok/s prefill, and full 262k context. Key settings: INT4 quant ideal for Ampere, weighted fp8 KV cache (near-identical to bf16), MTP speculative decoding (3 tokens), tensor parallel 2, prefix caching. Tuned over a month using vLLM metrics + Prometheus and vllm bench serve; full serve command posted. The author argues the overlooked RedHat quant is ideal for Ampere.
More from Infra
- Whittle distills on HF: 27B-A3B MoE quant claimed to run on 8GB VRAM laptops — depressedclassical · 2026-09-22
- Qdrant benchmark: post-upload latency spikes are optimizers, tuned configs yield up to 100x faster search — qdrant_engine · 2026-09-22
- Inference-free SPLADE: retrieval at BM25-like query cost without per-query inference — qdrant_engine · 2026-09-22
- What Cloudflare can't do: D1 caps at 10GB, is single-threaded, and no real Postgres — Paimaamu · 2026-09-22
- Rackspace joins NVIDIA Cloud Partner Program with Blackwell pods for regulated enterprises — DavidLinthicum · 2026-09-22
- Reverse-engineered Splash format ports Qwen3.8-27B to Mac, 88 tok/s at 64k on M5 Max — SeveralViolins · 2026-09-22