RTX 3090 runs Qwen3.8-27B at 672 tps via extreme quantization
iamMess · reddit · 2026-08-17
The author achieved peak speeds of 672 tps for Qwen3.6-28B on an RTX 3090 through extensive optimization. Key techniques include W4A16 quantization, FP8 KV cache, and int8 conversions for lmhead and embedtokens, reducing VRAM usage to 14.2GB. The setup achieves 82 tps for single requests and sustains 417 tps with 64 concurrent requests, outperforming ninfer by 17% to 149%. The solution runs on vLLM with a reported quality loss of approximately 0.6%.
More from Infra
- Bittensor Subnet 118 Adds Ultra-Cheap Inference, Joining Major AI Providers — markjeffrey · 2026-08-17
- Meta to rely on Nvidia Blackwell, AMD Helios in 2026, accelerate custom MTIA in 2027 — Beth_Kindig · 2026-08-17
- Stripe to Acquire OpenRouter for Over $7B, 5.4x May Valuation — rohanpaul_ai · 2026-08-17
- Wici One claims to solve local VRAM limits via NVMe offloading — Torodaddy · 2026-08-17
- Qwen3.8-27B hits 206 tok/s on single RTX 5090 via SGLang — StefanoGogioso · 2026-08-17
- antirez Optimizes DwarfStar: 170 t/s Generation and 22k tokens/s Prefill on Station — antirez · 2026-08-17