Running Kimi K3 (2.8T params) locally at full quality is possible

carrigmat · x · 2026-08-27

The author challenges the assumption that running the 2.8T-parameter Kimi K3 locally requires $100,000+ in hardware. They demonstrate a method to run the model at full quality locally, utilizing native experts and applying Q8 quantization strategies effectively.

Related event: Developer shows how to run full trillion-param LLMs locally on CPU for ~$6,000(14 posts)→

Original post →

More from Infra

Infra channel →