LithosAI launches ultra-fast inference with Kimi K3 at 800+ tokens/s
JiaZhihao · x · 2026-08-20
LithosAI announced its first pricing tiers, focusing on ultra-fast agentic inference on standard GPUs. The Kimi K3 model achieves over 800 tokens per second per user with full model quality. The Lithos Engine optimizes the serving system rather than degrading the model, supporting any agent harness like Claude Code or Codex. It offers both hosted API and on-prem deployments, with a public API launch set for September 1.
Related event: LithosAI Unveils Ultra-Fast Inference: Kimi K3 Hits 800+ Tokens/s(2 posts)→
More from Infra
- ffmpeg-webCLI: Local browser-based video editor — tom_doerr · 2026-08-20
- 75 data centers blocked or delayed by local opposition in Q1 — dinabass · 2026-08-20
- Monad Agent Hub launches with no-code platforms for instant agent creation — bgmshana · 2026-08-20
- Benchmark: MiniMax H3 Generation Speeds on AMD 9070 XT — Ok-Brain-5729 · 2026-08-20
- AWS Leverages AI Infrastructure Demand to Extend Cloud Dominance — DavidLinthicum · 2026-08-20
- Reverse-Engineering RK3588 NPU: Open Compiler Runs GPT-2 at 36 tok/s — one_does_not_just · 2026-08-20