vLLM Integrates TileRT, AMD MI355X Outpaces GB300 by 40% on GLM-5.3

vLLM and AMD used TileRT to run GLM-5.3 on 8×MI355X, reaching 469 tok/s single-user decode on the SemiAnalysis AgentX benchmark, 40% faster than GB300 FP4. The vLLM blog also details the TileRT integration, which decouples prefill and decode so the decode engine becomes pluggable for latency-sensitive serving.

2026-09-25 ~ 2026-09-25 · 2 related posts