Kimi K3 is reportedly running on 80 RTX 5090s with no HBM
markjeffrey · x · 2026-07-28
Kimi K3 is reportedly serving on 80 RTX 5090s with zero HBM
A quoted post says the full Kimi K3 model — described as 2.8 trillion parameters — is running on 80 RTX 5090s with 20 tok/s single-stream throughput on day one, without tuning.
Why it matters
The post claims this is a first for open weights: frontier-level intelligence served without HBM, using only GDDR7 gaming GPUs, standard Ethernet, and the official MXFP4 weights with no requantization.
The broader point is that the most powerful open model can now be hosted on relatively abundant consumer GPUs, which could make it easier for labs, startups, and universities to run agents, fine-tune, and probe the model themselves.
Related event: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(8 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23