vLLM Integrates TileRT, AMD MI355X Outpaces GB300 by 40% on GLM-5.3
vLLM and AMD used TileRT to run GLM-5.3 on 8×MI355X, reaching 469 tok/s single-user decode on the SemiAnalysis AgentX benchmark, 40% faster than GB300 FP4. The vLLM blog also details the TileRT integration, which decouples prefill and decode so the decode engine becomes pluggable for latency-sensitive serving.
2026-09-25 ~ 2026-09-25 · 2 related posts
- TileRT and AMD hit 469 tok/s decode on GLM-5.3 with vLLM on 8x MI355X, 40% faster than GB300 — vllm_project · 2026-09-25
- vLLM x TileRT: pluggable specialized decode for latency-critical serving, explained — vllm_project · 2026-09-25