LithosAI launches Kimi K3 with 800 tokens/sec inference speed on standard GPUs
Tim_Dettmers · x · 2026-08-20
LithosAI has announced its first pricing tiers and API launch, targeting ultra-fast agentic inference. The Kimi K3 model achieves over 800 tokens per second per user on standard GPUs while maintaining full model quality, claiming to be up to 16x faster than major providers.
Key Features:
- Extreme Speed: Tested at 800+ tokens/sec/user on a single 8×B300 node.
- Agent Optimized: Low latency across the entire agent loop (plan, call tools, observe, retry).
- High Compatibility: Offers OpenAI and Anthropic compatible APIs; works with Claude Code, Codex, and other harnesses by changing just the base URL.
- Model Support: Also serves GLM 5.3, Qwen 3.8, DeepSeek V4, and more.
Related event: LithosAI launches Kimi K3 with 800+ tokens/s inference(3 posts)→
More from Infra
- NVIDIA Releases CUDA-Q Algorithms, Open-Source Primitives for Fault-Tolerant Quantum Computing — tomaszbednarz · 2026-08-20
- Dolphin Network launches P2P inference network for idle GPUs — toptickcrypto · 2026-08-20
- Epoch estimate: OpenAI spends roughly 1.2–2.4% of compute on safety monitoring — sjgadler · 2026-08-20
- Hot take: Cerebras inference could speed up OpenAI's R&D loop 20x — Scobleizer · 2026-08-20
- AI Agents Are Not Microservices: The Need for Durable Execution — rseroter · 2026-08-20
- Using local Apple Intelligence models for in-app personalization — signulll · 2026-08-20