AMD MI355X Hits 952 tok/s Serving Kimi K3, Beating Nvidia B200/B300 in Efficiency
AnushElangovan · x · 2026-08-02
Engineers have successfully served the Kimi K3 model on AMD MI355X nodes, achieving remarkable inference performance.
- Aggregate Throughput: 952 tok/s/node
- Single Stream Decode: 118 tok/s
- Comparison: This outperforms the Nvidia B200 by 3.8x in aggregate throughput and 1.3x in single stream decode.
- Cost Efficiency: Beats the Nvidia B300 on performance per dollar (48 vs. 33 tok/s/$).
Related event: AMD MI355X Boosts Kimi K3 Throughput and Cuts Costs(3 posts)→
More from Infra
- AMD Enters Open-Source LLM Arena with Instella-MoE-16B — airesearch12 · 2026-08-03
- US States Move to Repeal Data Center Tax Breaks, Raising AI Infrastructure Costs — pstAsiatech · 2026-08-03
- Struggling with Local Video Models? Devs Discuss HuggingFace Pro ROI — Smooth-Telephone9443 · 2026-08-03
- Handling Offline AI Jobs: Developers Share Best Engineering Practices — cmm324 · 2026-08-03
- A 10-Week Roadmap for LLM Inference Serving and Optimization — _jaydeepkarale · 2026-08-03
- App Developers Should Ship Their Own On-Device Models — abacaj · 2026-08-03