GLM lists cache storage as 'limited-time free' while Kimi K3 charges $3.75 per cache write
teortaxesTex · x · 2026-09-21
A reposted thread flags a detail on Z AI's pricing page: both GLM-5.3 and 5.3 Flash mark cache storage as "limited-time free" — free for now, but leaving room to charge later. The author points to Kimi K3 on AWS Bedrock, which already charges $3.75 for cache writes (vs $0.30 read), more than regular input price — meaning the first cache-filling request costs more than sending tokens raw. If Z AI follows suit, GLM's effective cost could jump overnight without listed prices changing. Takeaway: when evaluating model pricing, the cache line is the one that can silently change the economics.
More from Models
- Jev's eval abstraction maps 1:1 to autorubric paper from 8 months ago, researcher finds — deliprao · 2026-09-21
- Mollick: the most annoying part of long agentic tasks is language drift, not hallucinations — emollick · 2026-09-21
- No, Laya isn't capped at 512 tokens — it's ModernBERT with 8192-token configs — antoine_chaffin · 2026-09-21
- The catapult analogy: why AI models are jagged and robots face a deployment gap — lateinteraction · 2026-09-21
- Five error modes of frontier models used raw at the API level — gerardsans · 2026-09-21
- Qwen Image 2.1 license draws flak: 'the worst license yet' — switch2stock · 2026-09-21