Kimi K3 runs on AMD MI350X with SGLang and hits 327 tok/s across four requests
burny_tech · x · 2026-07-28
Kimi K3 is shown running on AMD MI350X via SGLang, with the author reporting that it worked out of the box.
The attached benchmark screenshot suggests solid serving performance:
- Single-stream throughput ranges from about 105.6 tok/s on short prompts to 143.4 tok/s on reasoning prompts.
- In a 4-way parallel test, aggregate throughput reaches 327 tok/s across four requests.
The post is mainly a deployment note for AMD-based inference rather than a model announcement, and it implicitly signals growing compatibility between Kimi, SGLang, and AMD hardware.
Related event: SGLang Day-0 Support for Kimi K3 Hits 423 tok/s(5 posts)→
More from Infra
- Kimi K3 reportedly keeps block attention residuals in a 93-layer design — burny_tech · 2026-07-28
- llama.cpp benchmarks show ROCm and Vulkan trading wins on AMD Radeon AI PRO R9700 — Gesha24 · 2026-07-28
- NASA’s new administrator says SpaceX-backed orbital data centers will happen — elonmusk · 2026-07-28
- Reddit users test whether unlocked CMP 170HX cards can power local AI rigs — Tritheone69 · 2026-07-28
- Open-weight models are pitched as a security win for large companies — jessi_cata · 2026-07-28
- Dolphin is being trained on Trinity Large Thinking 398B across 72 RTX 4090s — QuixiAI · 2026-07-28