Qwen3.8-27B on 2x 3090 hits 218 tok/s decode with vLLM + DFlash2 spec-decode

xjx546 · reddit · 2026-08-19

A Reddit user benchmarks Qwen3.8-27B on 2× RTX 3090 (PCIe Gen4, no NVLink, patched P2P, power-capped 220/250W) using vLLM v0.26.1rc1 + AutoRound INT4 (group 128) + a DFlash2 draft model for speculative decoding:

Original post →

More from Infra

Infra channel →