Custom vLLM INT8 stack hits 972 tok/s on Qwen 27B with 4x MI100 ($6.5k rig)
1ncehost · reddit · 2026-08-27
A developer released a complete INT8 serving stack for older GPUs without native FP8: built on vLLM, AITER, and a GPTQ INT8 quant (DFlash2), it runs Qwen3.8 27B at 972 tok/s TG / 5,680 tok/s PP on a 4× MI100 rig ($6.5k), versus 15 tok/s on stock vLLM.
The work includes:
- A complete W8A8 INT8 GEMM library tuned for MI100, used everywhere
- INT8 KV cache and INT8 AITER Unified Attention (with Triton fallback, faster than Flash Attention)
- INT8 Mamba/GDN attention and INT8 embedding
- INT8 custom allreduce/allgather optimized for XGMI interlinks
- Many new fused INT8 kernels
Accuracy vetting is unusually rigorous: KLD measured not just per token but per GEMM, attention block, KV lookup, and layer, with diagnostic scripts left in the branch for verification. The author advises checking whether your card's INT8 TOPS exceeds its FP8 FLOPS — if so, this project applies; hardware-agnostic Triton fallbacks benefit MI50/MI210 and older cards too. Everything is open-sourced on GitHub and Hugging Face.
More from Infra
- RTX 5090 Config for Qwen 27B Local LLM: Params & Tips — Rollingsound514 · 2026-08-27
- Data centers' power-generation water use tops 3.4 trillion gallons a year in 7 states — AndyMasley · 2026-08-27
- Analyst: Nvidia Would Absorb All of Intel's Excess Capacity — BenBajarin · 2026-08-27
- Nvidia projects 70% revenue growth for fiscal 2028, beating analyst expectations of 44% — firstadopter · 2026-08-27
- Nvidia's NVHBM Brings 30% More Bandwidth to NVLink Fusion; Amazon Annapurna First Partner — nordicinst · 2026-08-27
- AWS and NVIDIA to Deploy 2 Million Additional GPUs for Agentic and Physical AI — nvidia · 2026-08-27