2.78T-param Kimi K3 runs inference on a single CPU in 8.24 GB of RAM
udmrzn · x · 2026-09-11
A demo shared via HuggingModels claims to run 2.78-trillion-parameter Kimi K3 inference on a single CPU with just 8.24 GB of RAM — no GPU, no BLAS, no framework. The approach implies aggressive weight streaming from disk, offering a radical path to local inference of frontier-scale models; accuracy and speed claims remain to be verified in the original project.
Related event: 2.78T-Parameter Kimi K3 Runs on Single CPU with Just 8.24GB RAM(3 posts)→
More from Infra
- Carmack: Jetson Thor's 128GB at 273GB/s is over-provisioned for real-time robotics — ID_AA_Carmack · 2026-09-11
- YC Demo Day startup touts ultra-pure diamond wafers for data centers, $160M in LOIs — ycombinator · 2026-09-11
- S2-Attention: hardware-aware Triton kernels make sparse attention actually fast — burkov · 2026-09-11
- Together launches preemptible compute for GPU clusters at 50% of on-demand price — togethercompute · 2026-09-11
- New Model Hits Opus-Level Benchmarks at Wild Efficiency, RL Infra Details Emerge — nrehiew_ · 2026-09-11
- SageMaker prefix-aware routing cuts P50 TTFT by up to 77% via warm KV caches — AWS ML Blog · 2026-09-11