Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM
The open-source Kimi K3 model (2.8 trillion parameters) has reportedly been successfully deployed on a cluster of 80 RTX 5090s, achieving 20 tok/s single-stream throughput with zero HBM. The setup uses official FP4 (MXFP4) weights without requantization, making it the first frontier LLM to run on pure consumer hardware, breaking the reliance on datacenter-grade hardware.
Confirmed
- Hardware: 80 RTX 5090s across 10 nodes (8 GPUs each), totaling 2.56TB VRAM.
- Network: 25GbE Ethernet between nodes, no InfiniBand.
- Model & Performance: Kimi K3 with 2.8T parameters, official FP4 weights, 20 tok/s single-stream.
- Self-hosting alternative: @Liueroteme proposed an EPYC 9556 + 16x 256GB DDR5 (4TB RAM) setup costing $50K; with 12800MT/s MRDIMM, bandwidth reaches 1.6TB/s, targeting 25 TPS.
Unconfirmed
- Posts mention that Kimi K3 and GLM-5.2 will soon open "Mining" mode; this remains a rumor.
Why it matters
- This is the first case of a frontier open-source LLM running on pure consumer hardware, breaking the absolute reliance on datacenter-grade hardware for high-end AI inference, offering a new path for low-cost deployment of ultra-large models.
- As top open-source model weights stabilize around 3T parameters, a $50K self-built server could become affordable for households. @Liueroteme compares it to buying a second car, suggesting that top-tier AI compute may become accessible in home settings.
2026-07-28 ~ 2026-07-29 · 8 related posts
- Episode 1: vLLM brings day-0 support to Moonshot’s Kimi K3(2026-07-27, 11 posts)
- Episode 2: Kimi K3 Now Available for Inference and Fine-Tuning on Fireworks(2026-07-28, 2 posts)
- Episode 3: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(2026-07-28, 8 posts)
- Episode 4: Kimi K3 Open-Weight Release Sparks Debate on Open Source and Infrastructure(2026-07-28, 5 posts)
- Episode 5: Kimi K3 Self-Hosting Can Break Even in Under 100 Days(2026-07-28, 4 posts)
- Episode 6: Tinkering with Local Quantized K3 Inference on Mac Hardware(2026-07-29, 2 posts)
- Episode 7: Kimi K3 Open Weights Demand Data Center Hardware(2026-07-29, 3 posts)
- Episode 8: vLLM Hits 464 tok/s on Kimi K3 with 4 GB300 Systems(2026-07-29, 2 posts)
- Episode 9: Kimi K3 Gets Day-0 vLLM and AMD Support Across Clouds(2026-07-30, 14 posts)
Primary sources
- Kimi K3 is reportedly running on 80 RTX 5090s with no HBM — markjeffrey ·
- Kimi K3 reportedly runs on 80 RTX 5090s with 2.56 TB of VRAM — sandyyevans ·
- K3 Model Self-Hosting: $50K Hardware Setup Achieves ~25 tps Inference — Liu_eroteme ·
- [source] Kimi K3 is reportedly running on 80 RTX 5090s with no HBM — markjeffrey · 2026-07-28
- Kimi K3 2.8T-parameter open model runs at 20 tok/s on 80 RTX 5090s with zero HBM — markjeffrey · 2026-07-28
- Kimi K3 reportedly runs on 80 RTX 5090s over 25GbE Ethernet — panchovix · 2026-07-28
- Kimi K3 and GLM-5.2 are said to open for mining soon on 80 RTX 5090s — const_reborn · 2026-07-28
- [source] Kimi K3 reportedly runs on 80 RTX 5090s with 2.56 TB of VRAM — sandyyevans · 2026-07-28
- [source] K3 Model Self-Hosting: $50K Hardware Setup Achieves ~25 tps Inference — Liu_eroteme · 2026-07-29
- Future Family Choice: Second Car or a K3 Model Server in the Basement? — Liu_eroteme · 2026-07-29
- A $50,000 EPYC self-hosted K3 build targets 1.6 TB/s bandwidth and ~25 TPS — Liu_eroteme · 2026-07-29