Running Kimi K3 on 8x B300: $190 per million tokens, full cost breakdown
OtherRaisin3426 · reddit · 2026-08-23
The author hosted the 2.8T-parameter Kimi K3 on Modal with 8x B300 ($56.79/hr), using vLLM with tensor parallel 8 and native MXFP4:
- Cold boot 27 min (1.56 TB load, JIT, 51 CUDA graph captures)
- TTFT 0.92–1.02 s, steady 92 tok/s decode, 83 tok/s average over 4 prompts
- $190 per million output tokens; a clean run costs $36 in GPU time; keeping it warm costs $1,363/day
As a comparison, Unsloth's 1-bit UD-IQ1S Dynamic GGUF (594 GB) fits 8x A100-80GB via llama.cpp at $19.99/hr (2.8x cheaper) — but delivers only 9 tok/s with 7–60 s TTFT, working out to $620 per million tokens (3.3x more expensive per token). Quality at 1-bit was fine (correct arithmetic, coherent prose). Full write-up with every flag, the Modal deployment file, and raw benchmark JSON on the linked blog.
More from Infra
- Mistral reportedly plans up to 1 GW of European compute capacity by 2030 — emmanuelvivier · 2026-08-23
- Contextual News Search APIs: A Deep Comparison for AI, RAG, and Research — ermanos12 · 2026-08-23
- Qwen3.8-27B MTP Grafted to Unsloth Saves RAM, Requires Thinking Mode — Nyghtbynger · 2026-08-23
- Nvidia AI Server Prices to Rise 15% Due to DRAM Shortage — The Decoder · 2026-08-23
- ComfyUI Node Optimization: Sparse Attention Boosts Speed by 5-20% — Zironic · 2026-08-23
- Benchmark Report: How Much Do Quants Matter on Modern Models? — KitchenAmoeba4438 · 2026-08-23