Why do we resend the whole conversation every turn? Server-side KV slots proposal sparks debate
Vasili_Sk · reddit · 2026-10-09
A Reddit post questions why LLM inference APIs force apps to resend the entire conversation each turn, making servers re-match KV slots and each harness reinvent compaction. The author proposes persistent server-side context slots — createslot/sendmessage/compact/setcachepolicy — with linear agent conversations skipping checkpoints entirely, and hot/cold tiering of KV managed server-side. Core argument: KV cache is the result of computation, not conversation history; inference servers should behave like databases (create→append→query→compact→delete) instead of reconstructing state from full replays. The author genuinely asks whether something fundamental blocks this design.
More from coding & agent
- Dev releases free 'AI Software Factory' ebook on agentic SDLC — Pavan_Belagatti · 2026-10-09
- Infisical Launches Agent Vault to Give AI Coding Agents API Access Without Real Credentials — ycombinator · 2026-10-09
- Forbes debuts Agentic 20 list spotlighting AI agents doing real-world work — annatonger · 2026-10-09
- Dev ships MCP server that lets Claude Code publish and share its tools — Zalaso · 2026-10-09
- No coding background: he shipped a Steam game in 9 months with Claude Code — Tommyruin · 2026-10-09
- Agent Plasticity: top-performing agents aren't the most efficient learners — RulinShao · 2026-10-09