Reddit user gets about 4 tokens/s running Kimi K3 on a 2×5090 home lab
iVoider · reddit · 2026-07-30
A Reddit post reports first home-lab results for Kimi K3: about 4 tokens per second on a machine with 768GB DDR5 and 2×5090 GPUs. The author also says prefill speed on long prompts reaches 50–70 tokens per second.
One oddity in the test is that decoding throughput appears to increase over time, which the author suspects may be due to warmup or swapping behavior. The benchmark tool crashed, so no formal comparison is shared.
More from Infra
- Baseten Launches Model Labs Platform for Closed-Model Monetization — baseten · 2026-07-30
- Baseten Launches Model Labs Platform for Commercializing Closed Models — baseten · 2026-07-30
- Replacing Cloud Vision APIs Locally with Nvidia Nemotron on DGX Spark — JFPuget · 2026-07-30
- Laguna XS Breaks 140 TPS on Apple Machines with New FAST Mode — gajesh · 2026-07-30
- $50B+ in AI Data Center Leases Signed in July as Bitcoin Miners Pivot — abhiadesai · 2026-07-30
- 26B-Parameter Gemma 4 Runs on Mac with 2GB RAM via SSD Streaming — petrusenko_max · 2026-07-30