Kimi K3 hits 800+ tokens/sec inference speed on standard GPUs

JiaZhihao · x · 2026-08-20

LithosAI announced that Kimi K3 is now running at 800+ tokens/sec/user on standard GPUs with full model quality. The team is pushing agentic inference to hardware limits and has released early-access pricing ahead of the September 1 API launch.

Related event: LithosAI Unveils Ultra-Fast Inference: Kimi K3 Hits 800+ Tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →