TileRT and AMD hit 469 tok/s decode on GLM-5.3 with vLLM on 8x MI355X, 40% faster than GB300
vllm_project · x · 2026-09-25
The vLLM project reports that TileRT and AMD teams pushed GLM-5.3 single-user decode to 469 tok/s on 8x MI355X in SemiAnalysis's AgentX benchmark — over 40% faster than GB300 TRTLLM on FP4. The setup is disaggregated: vLLM handles prefill while TileRT handles latency-critical decode via vLLM's V1 connector interface, keeping the rest of the deployment stock vLLM.
Related event: vLLM Integrates TileRT, AMD MI355X Outpaces GB300 by 40% on GLM-5.3(2 posts)→
More from Infra
- AMD to raise AI and consumer GPU prices ~10% in Q4 as TSMC hikes costs — Beth_Kindig · 2026-09-25
- Google's SunCatcher space data center satellite launches next week as compute leaves Earth — krishnan · 2026-09-25
- Wafer now powers Brilliant's AI tutor Koji, ending speculative prefetching for latency — ycombinator · 2026-09-25
- Pooling spare RAM across old devices to run 30B local models as memory prices soar — Medicine_Blogscanner · 2026-09-25
- 10-year Treasury tops 5.20% as traders say this cycle is compute-bound, not oil — km · 2026-09-25
- Smartphones eat ~30% of global DRAM and NAND supply — the fix? Stop yearly phone releases — AlpinDale · 2026-09-25