Keeping vLLM's prefix cache warm between agent turns
bolts98 · reddit · 2026-09-17
A hands-on engineering post on keeping vLLM's prefix cache (KV cache) warm across agent conversation turns, avoiding repeated prefill latency and compute waste, covering cache policy and server-side configuration practices.
Related event: Keeping vLLM Prefix Cache Hot Across Multi-Turn Agent Conversations(2 posts)→
More from coding & agent
- Open-source Helicon app brings Meta's Muse Code CLI to Windows desktop — alexandr_wang · 2026-09-17
- Netlify to livestream a Grok Bot autonomously building and deploying a site — thisiskp_ · 2026-09-17
- OpenJev open-sources Jev-style semantic decisions on a single RTX 3090 with a frozen 4B model — alexcovo_eth · 2026-09-17
- Browser Use CEO on agentic engineering: humans are the bottleneck, he stopped reading code — David Ondrej · 2026-09-17
- Jev: an open-source action-picker that splits agent thinking from clicking — alexcovo_eth · 2026-09-17
- CROA open-sources a deterministic execution layer enforcing trajectory-level constraints on AI agents — CROA_PROJECT · 2026-09-17