LLM Inference Costs Drop Below $3 with B200s, Yet API Prices Stay High

AccBalanced · x · 2026-07-30

A recent post highlights that serving large models is becoming significantly cheaper. Even with expensive B200 GPUs, vLLM, and without major optimizations, inference costs drop below $3. However, these savings aren't reflected in API endpoints, leaving users frustrated by the lack of affordable services like a $5 Kimi endpoint.

Original post →

More from Infra

Infra channel →