Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate

Vasili_Sk · reddit · 2026-10-09

A Reddit post questions why LLM inference APIs force apps to resend the entire conversation each turn, making servers re-match KV slots and each harness reinvent compaction. The author proposes persistent server-side context slots — createslot/sendmessage/compact/setcachepolicy — with linear agent conversations skipping checkpoints entirely, and hot/cold tiering of KV managed server-side. Core argument: KV cache is the result of computation, not conversation history; inference servers should behave like databases (create→append→query→compact→delete) instead of reconstructing state from full replays. The author genuinely asks whether something fundamental blocks this design.

Original post →

More from coding & agent

coding & agent channel →