Why long AI chats get expensive: full-history resends and brittle prefix caching

ClickOk5811 · reddit · 2026-10-10

The cost of a long chat session comes from resending the entire history each turn — a two-line reply on turn 40 carries 39 prior turns with it. Prompt caching only helps while the prefix stays byte-identical; edit an earlier message and everything after is a cache miss at full price. Bigger context windows make threads longer, not cheaper. The fix: start a new session when the task changes.

Original post →

More from Models

Models channel →