vLLM Achieves Major Throughput Gains on AMD GPUs
vLLM Blog · rss · 2026-07-13
A vLLM blog post details how to train, quantize, and deploy an EAGLE-3 speculative decoding draft model on AMD Instinct GPUs using vLLM + AMD Quark.
The core result is a significant boost in inference throughput, with reported figures including:
- Kimi-K2.5: Up to 2.00x throughput increase
- MiniMax-M2.5: Up to 1.79x throughput increase
The article focuses less on the models themselves and more on combining speculative decoding, quantization, and the serving stack to achieve higher serving efficiency on AMD hardware.
More from Infra
- Tabul AI launches Metal TreeSHAP to speed up Shapley values on Apple silicon — Scobleizer · 2026-07-22
- DeepSeek-V4-Flash tops out at 770 tok/s on one B300 in a vLLM batch test — Moreh · 2026-07-22
- NVIDIA starts shipping 102.4 Tbps Spectrum-6 switches for Vera Rubin AI factories — nvidia · 2026-07-22
- Apple publishes SOC 3 audit reports for Private Cloud Compute — throwfaraway4 · 2026-07-22
- Reddit GPU renters say existing platforms only give you two of three: code, recovery, fair billing — legendpizzasenpai · 2026-07-22
- The Sandboxing Manifesto: Secure Execution Environments for Agents — spirosoik · 2026-07-22