Post says Kimi K3 can cut compute by 60% and pay back self-hosting in under 100 days
JosephJacks_ · x · 2026-07-28
The post argues that Kimi K3 can cut more than 60% of compute by removing redundant matrix work and avoiding memory-bandwidth bottlenecks without changing output logits. It attributes much of the gain to Kimi Delta Attention, recurrent state caching, speculative decoding, Blackwell MXFP4 kernels, and asynchronous expert prefetching.
It also claims that, on 8× B300s, self-hosting K3 at about 75% utilization could serve over 30 billion tokens per month, with roughly $500K in compute plus under $10K in monthly opex. Using the cited API prices of $3 per million input tokens and $15 per million output tokens, the post estimates payback in under 100 days.
Related event: Kimi K3 Self-Hosting Can Break Even in Under 100 Days(4 posts)→
More from Infra
- Qualcomm goes agent-centric: Snapdragon 8 Elite Gen 6 and agent-native devices — jiqizhixin · 2026-09-23
- Unsloth Desktop Hotfix Adds Qwen-Image-2.1 Image Editing and Fixes GGUF Loading — danielhanchen · 2026-09-23
- Qwen 3.6 35B-A3B Q6 hits ~50 tok/s on a 128GB Strix Halo — what's the best local model now? — jankeydankey · 2026-09-23
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23