Kimi K3 is reportedly running on 80 RTX 5090s with no HBM
markjeffrey · x · 2026-07-28
Kimi K3 is reportedly serving on 80 RTX 5090s with zero HBM
A quoted post says the full Kimi K3 model — described as 2.8 trillion parameters — is running on 80 RTX 5090s with 20 tok/s single-stream throughput on day one, without tuning.
Why it matters
The post claims this is a first for open weights: frontier-level intelligence served without HBM, using only GDDR7 gaming GPUs, standard Ethernet, and the official MXFP4 weights with no requantization.
The broader point is that the most powerful open model can now be hosted on relatively abundant consumer GPUs, which could make it easier for labs, startups, and universities to run agents, fine-tune, and probe the model themselves.
Related event: Kimi K3 2.8T Model Runs Zero-HBM at 20 tok/s(2 posts)→
More from Infra
- Kimi K3 tokenizer optimization cuts first-token latency by about 325 ms — philipkiely · 2026-07-28
- Ben Bajarin says AI’s biggest miss was underestimating the GPU supply-chain tsunami — BenBajarin · 2026-07-28
- Palantir says its U.S. government AI platform is built on open weights and NVIDIA GPUs — eliano · 2026-07-28
- Kimi K3 lands day-one on Dell PowerEdge XE9780 for on-prem deployment — _akhaliq · 2026-07-28
- Claude Opus 5 leads benchmarks, but early users say it stops short on real work — The AI Daily Brief · 2026-07-28
- engy.ai launches verified inference API with cryptographic proof for open models like GLM-5.2 — markjeffrey · 2026-07-28