Kimi K3 inference on 8× AMD MI355X hits 216 tok/s per user with TokenSpeed
zhyncs42 · x · 2026-10-10
A team reports state-of-the-art Kimi K3 inference on 8× AMD MI355X GPUs using TokenSpeed: 216 output tok/s per user at C1, 1.4× ATOM's performance, at higher precision. The deep dive covers optimizations across KDA, MLA, AttnRes, LatentMoE, and multi-GPU communication, built on portable Triton with agent-assisted Gluon specialization.
Related event: TokenSpeed Hits New SOTA Running Kimi K3 on AMD MI355X(2 posts)→
More from Infra
- Cloudflare acquires Deno, will maintain runtime for only one more year — Simon Willison · 2026-10-10
- How apps scale: 2006 bigger servers, 2016 clusters, 2026 rewrite in Rust — tristanbob · 2026-10-10
- After HA Yellow failure and LLM-assisted eMMC debugging, altryne moves to Omarchy VM — altryne · 2026-10-10
- VidAIo claims AI video compression halves file size vs AWS, could cut Netflix's $1B streaming bill in half — markjeffrey · 2026-10-10
- Joseph Jacks: analog neural nets are going to be huge — your brain already runs them — JosephJacks_ · 2026-10-10
- Baseten launches Project Beacon, partners Goodfire for in-line open-model safety monitoring — baseten · 2026-10-10