Running Kimi K3 on CPU: Custom Q3 Quantization Takes 1.1TB, Hits 4.2 t/s

Fun-Meaning-6474 · reddit · 2026-07-30

A developer team used a custom fork of llama.cpp to quantize Kimi K3 (2.8T total parameters, 50B active) into GGUF format. The Q3KS version is currently working and takes up about 1114.76 GiB on disk.

Hardware & Performance:

Quality Testing:

The team tested the quantized model for text coherence and image understanding. For instance, they inputted a front page of the 1969 NYT moon landing, and the model accurately described the masthead, slogan, date, and headlines without hallucination. They are exploring whether running such massive quantized models on CPUs is a viable alternative to smaller, faster models.

Original post →

More from Infra

Infra channel →