Running Kimi K3 on a B300: 450 tokens/s for $46k/month
casper_hansen_ · x · 2026-08-01
A developer ran the economics on self-hosting: for heavy consumers burning billions of tokens daily, renting a server with a single B300 GPU for $46,000 a month to run Kimi K3 could be a highly viable alternative to Claude. This setup also provides exclusive access to blazing-fast inference speeds of 450 tokens/s.
More from Infra
- Atomic-Chat: An Open-Source Local AI Assistant That Runs 100% Offline — rohanpaul_ai · 2026-08-01
- 1-bit Kimi K3 Quant Tested: 2.8T Model Compressed to 590GB Runs Locally — rohanpaul_ai · 2026-08-01
- Paper Share: How Chunked Prefill Improves LLM Serving Efficiency — Abhishekcur · 2026-08-01
- The Compute Bottlenecks of Agentic AI: Inference vs. Execution — charles_irl · 2026-08-01
- AI Market Correction Warning: Extreme Leverage in Memory Chips and Record Investor Debt — binarybits · 2026-08-01
- DeepSeek-V4-Flash Quantized on A100: Uses Only 15.8GB VRAM at 16 tok/s — Different-Pickle1021 · 2026-08-01