Running 1.6TB Kimi K3 Weights: 128GB Mac vs 80x RTX 5090 Cluster
机器之心 · wechat · 2026-08-01
After the 2.8 trillion parameter Kimi K3 opened its weights, local deployment became a challenge. A developer successfully loaded the 1.56TB official MXFP4 weights on a 128GB unified memory Mac by continuously streaming from the SSD. Although the model functioned, the generation speed was only 0.32 token/s.
Another team solved the performance issue by stacking consumer GPUs. They used 80 RTX 5090s across 10 nodes, providing 2.56TB of total VRAM. Running the unquantified official Kimi K3 weights, the cluster achieved a first-day inference speed of 20 token/s, proving that consumer-grade GPU clusters can run ultra-large models at usable speeds even without HBM or high-speed interconnects.
More from Infra
- SDNQ Quantization Engine Integrated into Diffusers with Multi-Platform Support — RisingSayak · 2026-08-01
- NXP Semiconductors in Talks to Acquire AI Chip Designer Ambarella — pstAsiatech · 2026-08-01
- Full 2.78T-parameter Kimi K3 Runs on Consumer Laptop via NVMe Streaming — rickasaurus · 2026-08-01
- CXMT's LPDDR6 Memory Nearing Mass Production with 12,800Mbps Speed — bookwormengr · 2026-08-01
- OpenAI Hits Git Perf Limits in Giant Monorepo, Upstreams Fixes — charliermarsh · 2026-08-01
- Why Chinese LLMs Struggle in AI Coding: The Hidden Costs of Compute and Quotas — 创业邦 · 2026-08-01