Kimi K3 2.8T-parameter open model runs at 20 tok/s on 80 RTX 5090s with zero HBM
markjeffrey · x · 2026-07-28
The Kimi K3 model with 2.8 trillion parameters achieves 20 tok/s single-stream inference on 80 RTX 5090 GPUs, requiring no HBM memory, only GDDR7 gaming cards and plain ethernet. It uses official MXFP4 weights without requantization, marking the first frontier open model running on consumer GPUs. Any lab, startup, or university can own, probe, fine-tune, and run agents on it.
Related event: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(8 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23