vLLM Achieves Major Throughput Gains on AMD GPUs
vLLM Blog · rss · 2026-07-13
A vLLM blog post details how to train, quantize, and deploy an EAGLE-3 speculative decoding draft model on AMD Instinct GPUs using vLLM + AMD Quark.
The core result is a significant boost in inference throughput, with reported figures including:
- Kimi-K2.5: Up to 2.00x throughput increase
- MiniMax-M2.5: Up to 1.79x throughput increase
The article focuses less on the models themselves and more on combining speculative decoding, quantization, and the serving stack to achieve higher serving efficiency on AMD hardware.
More from Infra
- SF Compute founder: buying compute is 'an absolutely awful experience' right now — IgorCarron · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- PyTorch Day Korea 2026 launches first offline conf, CFP closes Sept 13 — PyTorch · 2026-09-11
- Local LLM server dilemma: 4x CMP-170HX (price up 53% in 20 days) vs Mac Studio M5 Ultra — rumboll · 2026-09-11
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Your p99 latency benchmark may be lying: a deep dive into coordinated omission — Franc0Fernand0 · 2026-09-11