Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM

The open-source Kimi K3 model (2.8 trillion parameters) has reportedly been successfully deployed on a cluster of 80 RTX 5090s, achieving 20 tok/s single-stream throughput with zero HBM. The setup uses official FP4 (MXFP4) weights without requantization, making it the first frontier LLM to run on pure consumer hardware, breaking the reliance on datacenter-grade hardware.

Confirmed

Unconfirmed

Why it matters

2026-07-28 ~ 2026-07-29 · 8 related posts

Full story(9 episodes)→

Primary sources