Full 2.78T-parameter Kimi K3 Runs on Consumer Laptop via NVMe Streaming
rickasaurus · x · 2026-08-01
Developer Marco Bambini introduced WASTE (Weight-Aware Streaming Tensor Engine), a clean C-based inference engine that successfully runs the unmodified, 2.78-trillion-parameter Kimi K3 model on a consumer laptop.
The core of this approach is weight-aware streaming: it bypasses VRAM limits by streaming only the activated experts directly from NVMe storage. The process requires no distillation, pruning, or cloud connection, preserving the full open weights. While currently slow, it proves the technical feasibility of running massive MoE models on local hardware.
More from Infra
- SDNQ Quantization Engine Integrated into Diffusers with Multi-Platform Support — RisingSayak · 2026-08-01
- Running 1.6TB Kimi K3 Weights: 128GB Mac vs 80x RTX 5090 Cluster — 机器之心 · 2026-08-01
- NXP Semiconductors in Talks to Acquire AI Chip Designer Ambarella — pstAsiatech · 2026-08-01
- CXMT's LPDDR6 Memory Nearing Mass Production with 12,800Mbps Speed — bookwormengr · 2026-08-01
- OpenAI Hits Git Perf Limits in Giant Monorepo, Upstreams Fixes — charliermarsh · 2026-08-01
- Why Chinese LLMs Struggle in AI Coding: The Hidden Costs of Compute and Quotas — 创业邦 · 2026-08-01