Users dispute credit burn; provider says KV cache was always on, scaling across providers
arthurcolle · x · 2026-09-04
A user publicly disputed a provider's billing: massive credits spent in the first days after release made them suspect KV cache wasn't enabled; they also argued subscription plans are poor value since API discounts aren't what subscriptions should target. The provider replied that KV cache has always been enabled and that managing scaling across providers and increasing cache hits is its first priority. The exchange exposes real tensions between cache economics and billing transparency.
Related event: Abliteration denies KV cache disabled amid billing complaints(2 posts)→
More from Infra
- Inference startup insider: "we just resell NVIDIA GPUs" — VCs question the moat — firstadopter · 2026-09-04
- Leak claims GPT-6 Astra trained on 100,000+ GPUs at OpenAI's Stargate site — BLUECOW009 · 2026-09-04
- Modal adds support for running Cursor Cloud Agents in custom sandboxes — AAAzzam · 2026-09-04
- One cluster alone could train 68 GPT-6-scale models by 2029, and FP4 could double that — scaling01 · 2026-09-04
- NVIDIA Jetson initrd flaw lets attackers with physical access bypass Secure Boot — jedisct1 · 2026-09-04
- Ollama's new interactive menu makes launching local models and agents easier — Technovangelist · 2026-09-04