Qwen3.8-27B on 2x 3090 hits 218 tok/s decode with vLLM + DFlash2 spec-decode
xjx546 · reddit · 2026-08-19
A Reddit user benchmarks Qwen3.8-27B on 2× RTX 3090 (PCIe Gen4, no NVLink, patched P2P, power-capped 220/250W) using vLLM v0.26.1rc1 + AutoRound INT4 (group 128) + a DFlash2 draft model for speculative decoding:
- Code decode hits 218.3 tok/s, narrative 120.1 tok/s; TTFT 170ms
- Prefill: 1342 tok/s @ 10k context, 628 tok/s @ 90k
- Spec-decode: 7 draft tokens, 47.8% acceptance, avg accepted length 3.35
- Peak VRAM 22.3 GB/card with zero leak; context ceiling 131k (drafter uses 13.5GB)
- Measured with the Club-3090 canonical bench suite (3 warmups + 5 runs, temp 0.6/topp 0.95/topk 20); author says more performance is on the table
- Custom vLLM changes were needed to boot; all vLLM fixes were done with Kimi K3
More from Infra
- DFlash 2 released: up to 4.6× speedup for AI inference — igilitschenski · 2026-08-19
- Will we run 30B+ parameter models fast on small GPUs in the future? — absurdother · 2026-08-19
- Periodic Labs trains trillion-parameter models on Miles framework, 3x throughput boost — hsu_byron · 2026-08-19
- Local LLM Speed Bottlenecks: RTX 4090 vs. 5090 Performance Analysis — Viktri1 · 2026-08-19
- LLM Inference Engineering: From KV Cache to vLLM and SGLang — techNmak · 2026-08-19
- NVIDIA H100 Concurrency Response of Plain Global Loads Analyzed — ssh4net · 2026-08-19