Kimi-K3 runs locally at 0.23 tok/s on a dual RTX 6000 PRO workstation
Aroochacha · reddit · 2026-07-29
A user reports getting Kimi-K3 running locally through a llama.cpp PR and shares very slow throughput numbers.
- Prompt evaluation: 40 tokens in 97.5s (0.41 tok/s)
- Main eval: 400 tokens in 1769.9s (0.23 tok/s)
- Total: 440 tokens in 1867s (31 minutes)
- The setup used a high-end workstation with dual RTX 6000 PRO 96GB cards and large NVMe RAID storage.
- The post also notes the next step is wiring the workstation into a 100GbE fabric for RPC use.
More from coding & agent
- Tencent Releases UI-Mate-27B, a Desktop GUI Agent Model — tencent · 2026-08-24
- Comparing AI Subscriptions: DeepSeek API vs. Claude Pro vs. Local LLMs — Unlikely_Bluejay5392 · 2026-08-24
- Claude Code introduces 'Remote Control' feature to boost coding efficiency — rohanpaul_ai · 2026-08-24
- rauchg lays out fx extension philosophy: MCP, Skills, Plugins and Unix composition — AccBalanced · 2026-08-24
- Netflix details its production LLM judge: hundreds of thousands of recommendations scored weekly — omarsar0 · 2026-08-24
- smolvm passes Simon Willison's Fable 5 agent test as a secure sandbox — yawnxyz · 2026-08-24