Kimi K3 2.8T-parameter open model runs at 20 tok/s on 80 RTX 5090s with zero HBM
markjeffrey · x · 2026-07-28
The Kimi K3 model with 2.8 trillion parameters achieves 20 tok/s single-stream inference on 80 RTX 5090 GPUs, requiring no HBM memory, only GDDR7 gaming cards and plain ethernet. It uses official MXFP4 weights without requantization, marking the first frontier open model running on consumer GPUs. Any lab, startup, or university can own, probe, fine-tune, and run agents on it.
Related event: Kimi K3 2.8T Model Runs Zero-HBM at 20 tok/s(2 posts)→
More from Infra
- Kimi K3 tokenizer optimization cuts first-token latency by about 325 ms — philipkiely · 2026-07-28
- Ben Bajarin says AI’s biggest miss was underestimating the GPU supply-chain tsunami — BenBajarin · 2026-07-28
- Palantir says its U.S. government AI platform is built on open weights and NVIDIA GPUs — eliano · 2026-07-28
- Kimi K3 lands day-one on Dell PowerEdge XE9780 for on-prem deployment — _akhaliq · 2026-07-28
- Claude Opus 5 leads benchmarks, but early users say it stops short on real work — The AI Daily Brief · 2026-07-28
- engy.ai launches verified inference API with cryptographic proof for open models like GLM-5.2 — markjeffrey · 2026-07-28