Kimi Code Model Inference Exceeds 1000 tok/s

JiaZhihao · x · 2026-07-10

Lithos shared the serving stack they built for Kimi K2.7 Code: on a single 8×B200 node, it achieves a peak throughput of over 1000 tokens/s per user for coding workloads, maintaining native model precision and full quality without introducing extra approximation.

They emphasized that for agents, speed directly impacts the latency of every reasoning, coding, and iteration step; thus, "ultra-fast inference" will fundamentally transform the agentic experience. The post includes a trial link.

Original post →

More from coding & agent

coding & agent channel →