TokenSpeed adds day-0 Kimi K3 support on NVIDIA Blackwell and AMD MI350X chips
zhyncs42 · x · 2026-07-27
TokenSpeed brings Kimi K3 to NVIDIA Blackwell and AMD MI350X/MI355X on day 0
TokenSpeed says it worked with Moonshot AI to add day-0 support for Kimi K3 on NVIDIA (G)B200/(G)B300 and AMD Instinct MI350X/MI355X within a week of the model announcement.
The post highlights the serving stack used to make that happen:
- Prefix caching
- Speculative decoding
- Disaggregated serving
- CUDA Graph decode
It also says the team optimized KDA, Gated MLA, Stable LatentMoE, AttnRes, and a unified Flat KV architecture, while laying groundwork for next-gen platforms like NVIDIA Vera Rubin and AMD MI455X.
The blog frames Kimi K3 as a 2.8T-parameter MoE transformer with a 1M-token context window, and claims it is a frontier model with strong long-horizon reasoning and coding performance. It also cites benchmark results including AutomationBench 30.8, SpreadsheetBench 2 34.8, BrowseComp 91.2, and DeepSWE 67.5.
More from Infra
- Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving — ricklamers · 2026-07-27
- How to run a private local AI on many 16GB laptops with Ollama and Gemma 4 — minchoi · 2026-07-27
- Vercel adds Kimi K3 on US providers with zero-data-retention support — cramforce · 2026-07-27
- Kimi K3 is said to cost more than 2× as much to serve as V4 — teortaxesTex · 2026-07-27
- Cisco releases Antares, open-weight models for local code-vulnerability hunting — thione · 2026-07-27
- Etched raises $300M at a $10.3B valuation to push its AI chip strategy — thione · 2026-07-27