TileRT and AMD hit 469 tok/s decode on GLM-5.3 with vLLM on 8x MI355X, 40% faster than GB300

vllm_project · x · 2026-09-25

The vLLM project reports that TileRT and AMD teams pushed GLM-5.3 single-user decode to 469 tok/s on 8x MI355X in SemiAnalysis's AgentX benchmark — over 40% faster than GB300 TRTLLM on FP4. The setup is disaggregated: vLLM handles prefill while TileRT handles latency-critical decode via vLLM's V1 connector interface, keeping the rest of the deployment stock vLLM.

Related event: vLLM Integrates TileRT, AMD MI355X Outpaces GB300 by 40% on GLM-5.3(2 posts)→

Original post →

More from Infra

Infra channel →