Post says Kimi K3 can cut compute by 60% and pay back self-hosting in under 100 days
JosephJacks_ · x · 2026-07-28
The post argues that Kimi K3 can cut more than 60% of compute by removing redundant matrix work and avoiding memory-bandwidth bottlenecks without changing output logits. It attributes much of the gain to Kimi Delta Attention, recurrent state caching, speculative decoding, Blackwell MXFP4 kernels, and asynchronous expert prefetching.
It also claims that, on 8× B300s, self-hosting K3 at about 75% utilization could serve over 30 billion tokens per month, with roughly $500K in compute plus under $10K in monthly opex. Using the cited API prices of $3 per million input tokens and $15 per million output tokens, the post estimates payback in under 100 days.
Related event: Kimi K3 Self-Hosting Saves 60% Compute, Pays Off in 100 Days(2 posts)→
More from Infra
- Bloomberg says Nvidia has $750B of AI deals, stoking worries about circular demand — FinanceYF5 · 2026-07-28
- Kimi K3 starts serving in Japan with day-one inference on B300 and MI355X — DavidBennett__ · 2026-07-28
- Nvidia reportedly weighs $250B OpenAI backstop while backing SSI with a $5B bet — 新智元 · 2026-07-28
- Researcher weighs local hardware after RunPod costs keep rising on small LLM runs — lupodevelop · 2026-07-28
- User asks whether ROCm 7.14 is worth upgrading for llama.cpp inference on Ubuntu — DeepBlue96 · 2026-07-28
- A Stanford talk highlights a startup that cut its AI bill from $1.2M to about $100k a month — HeyAmit_ · 2026-07-28