LithosAI launches ultra-fast inference with Kimi K3 at 800+ tokens/s

JiaZhihao · x · 2026-08-20

LithosAI announced its first pricing tiers, focusing on ultra-fast agentic inference on standard GPUs. The Kimi K3 model achieves over 800 tokens per second per user with full model quality. The Lithos Engine optimizes the serving system rather than degrading the model, supporting any agent harness like Claude Code or Codex. It offers both hosted API and on-prem deployments, with a public API launch set for September 1.

Related event: LithosAI Unveils Ultra-Fast Inference: Kimi K3 Hits 800+ Tokens/s(2 posts)→

Original post →

More from Infra

Infra channel →