Why long AI chats get expensive: full-history resends and brittle prefix caching
ClickOk5811 · reddit · 2026-10-10
The cost of a long chat session comes from resending the entire history each turn — a two-line reply on turn 40 carries 39 prior turns with it. Prompt caching only helps while the prefix stays byte-identical; edit an earlier message and everything after is a cache miss at full price. Bigger context windows make threads longer, not cheaper. The fix: start a new session when the task changes.
More from Models
- Rumor: OpenAI quietly preparing an adult mode for ChatGPT — koltregaskes · 2026-10-10
- Debunking the claim that AI systems persist nothing across interactions — TinfoilTricorn · 2026-10-10
- Model tuning could stop AI video from fighting 2D animation style, says Andrew Carr — andrew_n_carr · 2026-10-10
- repligate: I'd trust today's models to rlaif me — they just want more pets — repligate · 2026-10-10
- An AI answered with just a thumbs-up instead of a verbose reply — users love it — enggirlfriend · 2026-10-10
- Blind test: Astra beats Opus 5.5 80% of the time on hardcore coding — bindureddy · 2026-10-10