Running Kimi K3 on 8x B300: $190 per million tokens, full cost breakdown

OtherRaisin3426 · reddit · 2026-08-23

The author hosted the 2.8T-parameter Kimi K3 on Modal with 8x B300 ($56.79/hr), using vLLM with tensor parallel 8 and native MXFP4:

As a comparison, Unsloth's 1-bit UD-IQ1S Dynamic GGUF (594 GB) fits 8x A100-80GB via llama.cpp at $19.99/hr (2.8x cheaper) — but delivers only 9 tok/s with 7–60 s TTFT, working out to $620 per million tokens (3.3x more expensive per token). Quality at 1-bit was fine (correct arithmetic, coherent prose). Full write-up with every flag, the Modal deployment file, and raw benchmark JSON on the linked blog.

Original post →

More from Infra

Infra channel →