RTX 3090 runs Qwen3.8-27B at 672 tps via extreme quantization

iamMess · reddit · 2026-08-17

The author achieved peak speeds of 672 tps for Qwen3.6-28B on an RTX 3090 through extensive optimization. Key techniques include W4A16 quantization, FP8 KV cache, and int8 conversions for lmhead and embedtokens, reducing VRAM usage to 14.2GB. The setup achieves 82 tps for single requests and sustains 417 tps with 64 concurrent requests, outperforming ninfer by 17% to 149%. The solution runs on vLLM with a reported quality loss of approximately 0.6%.

Original post →

More from Infra

Infra channel →