vLLM details MiniMax M3 optimization on AMD MI355X: 4.45x per-GPU throughput gains
vllm_project · x · 2026-09-11
The vLLM team, with AMD and EmbeddedLLM, published a deep dive on optimizing MiniMax M3 on Instinct MI355X after day-0 support, following the bottleneck as it moves.
Key results from the SemiAnalysis InferenceX benchmark:
- MXFP8 standard serving at concurrency 32 rose from 109.1 to 342.4 output tokens/s/GPU (3.14x); median TTFT fell from 1.46s to 0.67s
- At concurrency 128, the TP4/EP1 path reached 623.7 (2.09x)
- MXFP4 with TP2/EP1 hit 943.5 tokens/s/GPU — 4.45x the initial per-GPU result
- EAGLE3 speculative decoding reached 682.4; P/D disaggregation delivered 6,370.5 total tokens/s/GPU at concurrency 512
The post covers MSA sparse attention, quantization, topology tuning, and reusable performance tips for future model optimizations.
More from Infra
- Frontier models now independently reach for speculative decoding and kernel optimization on InferenceBench — maksym_andr · 2026-09-11
- A Beginner-Friendly Guide to Budget Multi-GPU Local LLM Setups — lblblllb · 2026-09-11
- Chinese Nvidia challenger Enflame jumps 179% in Shanghai debut, raises $910M — pstAsiatech · 2026-09-11
- Qwen3.8 Flash Next hits 49 tok/s locally on 2x RTX 3090 with FlashNext llama.cpp fork — whiteh4cker · 2026-09-11
- Qdrant lines up three free community events with 4-hour vector tech stream — qdrant_engine · 2026-09-11
- B300 spot prices hit $2.2M per unit in China, 3x premium pushes domestic chips into the推理 sweet spot — aigclink · 2026-09-11