176 KB C Program Runs 2.78T-Param Kimi K3 on a Single CPU with 8.24 GB RAM
techNmak · x · 2026-09-11
Someone wrote a 176 KB C program that runs the 2.78 trillion-parameter Kimi K3 on a single CPU with 8.24 GB of RAM — and the author of this thread dug into the repo to verify how. The checkpoint isn't shrunk: it's still 1.56 TB.
- MoE sparsity is the trick: 92 routed layers with 896 experts each, only 16 active per token — 82,432 routed experts occupying 1.447 TB (93% of the checkpoint), parked on NVMe instead of RAM
- Each expert is 17.56 MB; a token can require 1,472 expert fetches across layers, with the engine reading 25.83 GB of expert weights per token without resident hits
- The always-used 113 GB of weights are repacked to fit the tight memory budget
A striking example of combining MoE sparsity with NVMe streaming for "poor man's" large-model inference.
Related event: 2.78T-parameter Kimi K3 runs on a single CPU with 8.24GB RAM(4 posts)→
More from Infra
- PlanetScale's Neki sharded Postgres hits 118M QPS across 512 shards — DanielLockyer · 2026-09-12
- Who Actually Pays Together, Fireworks and DeepInfra? A Reddit Debate — Azamat_Kuzdibay · 2026-09-12
- Reddit Petitions llama.cpp for Hot Expert Reload to Speed Up Local MoE Inference — perelmanych · 2026-09-12
- Together AI expands fine-tuning with GLM-5.3, Kimi K2.7-Code, live metrics, 30-70% price cuts — togethercompute · 2026-09-12
- 31 million protein complex predictions run on NVIDIA BioNeMo, saving an estimated 1.35 GWh — AllThingsApx · 2026-09-12
- Signal65 launches PINNACLE, an agentic AI benchmark scoring correct work over raw throughput — ryanshrout · 2026-09-12