Kimi K3 hits 800+ tokens/sec inference speed on standard GPUs
JiaZhihao · x · 2026-08-20
LithosAI announced that Kimi K3 is now running at 800+ tokens/sec/user on standard GPUs with full model quality. The team is pushing agentic inference to hardware limits and has released early-access pricing ahead of the September 1 API launch.
Related event: LithosAI Unveils Ultra-Fast Inference: Kimi K3 Hits 800+ Tokens/s(2 posts)→
More from Infra
- 1.5-Year Delay Cuts AI Data Center Value by 8.9%, Speed Key: Report — Beth_Kindig · 2026-08-20
- Clarification: OpenAI's 20% compute claim refers to monitoring overhead, not total capacity — sjgadler · 2026-08-20
- SkyPilot, VAST Data, and Partners Host AI Infra Meetup — skypilot_org · 2026-08-20
- PolymathicAI Releases The Well: A 15TB Collection of Physics Simulations — tom_doerr · 2026-08-20
- Filesystems beat MCP calls by 100x for context retrieval — ml_guy1 · 2026-08-20
- Addy Osmani on Engineering Roles and AI Agents — addyosmani · 2026-08-20