Batching vLLM calls with shared prefixes: a warm-up-then-fan-out KV cache strategy
Theboyscampus · reddit · 2026-09-30
A developer describes an agent workflow running 10 concurrent vLLM API calls via asyncio.gather that share a common prompt prefix, routed through a vllm-router/llm-d KV-cache-aware router to a pool of vLLM workers. Their proposed best practice: send one request first so vLLM completes a cache block and starts decoding, then fan out the rest of the batch, with the router pinning all requests to the same worker to maximize prefix cache hits. The thread solicits feedback on this warm-up-then-fan-out approach.
More from coding & agent
- 5 ways to cut LLM costs without changing models: optimize tokens, caching and calls — goyalshaliniuk · 2026-09-30
- Open Pstack: Codex as parent, Devin implements, Claude reviews — multi-provider agent orchestration — CombinationOk2374 · 2026-09-30
- VFX creator uses Opus 5.5 to drive Blender, freezing dancers into statues — anselm · 2026-09-30
- Cisco's 2026 campus hiring adds AI-agent project round with prompt documentation — cneuralnetwork · 2026-09-30
- One Year of Local Image Generation: Why Civitai and ComfyUI Both Fall Short — BenDLH · 2026-09-30
- SpatialClaw: Training-Free Code-Action Agent Beats Prior by 11.2 Points on 20 Spatial Benchmarks — _akhaliq · 2026-09-30