Batching vLLM calls with shared prefixes: a warm-up-then-fan-out KV cache strategy

Theboyscampus · reddit · 2026-09-30

A developer describes an agent workflow running 10 concurrent vLLM API calls via asyncio.gather that share a common prompt prefix, routed through a vllm-router/llm-d KV-cache-aware router to a pool of vLLM workers. Their proposed best practice: send one request first so vLLM completes a cache block and starts decoding, then fan out the rest of the batch, with the router pinning all requests to the same worker to maximize prefix cache hits. The thread solicits feedback on this warm-up-then-fan-out approach.

Original post →

More from coding & agent

coding & agent channel →