AMD MI355X runs Kimi K3 on single node, 3.8x throughput of B200 setup, better cost-performance than B300
机器之心 · wechat · 2026-08-04
Jiqizhixin reports that WaferAI deployed Kimi K3 on AMD MI355X, requiring only 8 MI355X GPUs in one server, while B200 needed 16 GPUs across two servers. Measured total throughput reached 952 tokens/s, single-user generation speed 118 tokens/s, about 3.8x the per-node throughput of the 16-GPU B200 setup, with better cost-performance than B200 and B300. ROCm support was smooth with minor fixes. Initial TTFT was slow but improved significantly after padding attention heads. The article highlights that memory capacity becomes critical for large models, and AMD's large HBM strategy is becoming a system advantage.
More from Infra
- NVIDIA Sol Engine Accelerates MiniMax H3 by 3.95x End-to-End — KissMyShinyArse · 2026-08-04
- AI Shipping Labs Launches 'Inference Engineering' Book Club on Aug 10 — Al_Grigor · 2026-08-04
- FMS 2026 Kicks Off: Samsung, SK Hynix, Nvidia to Discuss Memory-Centric AI Vision — AccBalanced · 2026-08-04
- AMD Faces AI Earnings Test as Open Models Threaten Closed Labs' Margins — AccBalanced · 2026-08-04
- RTX 5090 Hits 83°C Running MiniMax H3 Locally — Careless-Constant-33 · 2026-08-04
- MiniMax H3 Text-to-Video Successfully Runs on a Single DGX Spark — Scobleizer · 2026-08-04