Kimi K3 reportedly runs on 8GB RAM, but needs 33 seconds per token
petrusenko_max · x · 2026-08-04
Kimi K3 is reported to run on just 8GB of RAM, but at a comically slow pace of 33 seconds per token.
The post notes that the setup activates only 16 experts at a time, framing the result as a joke: not 33 tokens per second, but 33 seconds per token.
More from Models
- xAI updates Grok Build with Grok 4.5, skills, MCP, and plan mode — elonmusk · 2026-08-04
- RL on custom search harnesses may beat the “one big model” idea — shangbinfeng · 2026-08-04
- Qwen 3.8 Max reaches 42% on the hard INDUCTION benchmark, taking second place — DeryaTR_ · 2026-08-04
- OpenAI hires the creator of WebRTC as its GPT-Live voice system gets a deep dive — bookwormengr · 2026-08-04
- Code Arena WebDev puts four open-weight Chinese models near the frontier — floriandotorg · 2026-08-04
- GPT-5.6 Misspells Email Address, Then Hallucinates a Post-Hoc Excuse — WolframRvnwlf · 2026-08-04