AMD MI355X runs Kimi K3 on single node, 3.8x throughput of B200 setup, better cost-performance than B300

机器之心 · wechat · 2026-08-04

Jiqizhixin reports that WaferAI deployed Kimi K3 on AMD MI355X, requiring only 8 MI355X GPUs in one server, while B200 needed 16 GPUs across two servers. Measured total throughput reached 952 tokens/s, single-user generation speed 118 tokens/s, about 3.8x the per-node throughput of the 16-GPU B200 setup, with better cost-performance than B200 and B300. ROCm support was smooth with minor fixes. Initial TTFT was slow but improved significantly after padding attention heads. The article highlights that memory capacity becomes critical for large models, and AMD's large HBM strategy is becoming a system advantage.

Original post →

More from Infra

Infra channel →