Kimi K3 and GLM-5.2 are said to open for mining soon on 80 RTX 5090s
const_reborn · x · 2026-07-28
Mining for GLM-5.2 and Kimi K3 is said to open soon.
The quoted post says Kimi K3 was run as a full open-weights model on 80 RTX 5090s, reaching 20 tok/s on a single stream on day one, without tuning. It also claims GLM-5.2 was improved from 30 tok/s to 110 tok/s on the same fleet last week, and that performance should keep climbing.
The key takeaway is the infrastructure angle: the author frames this as frontier-scale inference with no HBM, using only GDDR7 gaming cards, plain Ethernet, and official MXFP4 weights — meaning labs, startups, and universities could own, inspect, fine-tune, and run agents on it.
Related event: Kimi K3 2.8T-Parameter Model Runs on 80 RTX 5090s with Zero HBM(8 posts)→
More from Infra
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23