Workload-aware inference: why batch LLM pipelines should plan queries like databases do

sh_reya · x · 2026-10-02

The author is bullish on workload-aware inference, arguing database systems are the crown jewel of exploiting workload knowledge for performance—and any system with a declarative interface and budding LLM inference demand is a good candidate.

In a batch data processing pipeline with multiple LLM operations, the typical API approach makes one call per row per operation, served independently. If you know all requests up front, you can plan the whole job: reorder requests within an operation to share KV cache, run the most selective filter first, and reuse each document's KV cache across operations. Beyond planning, execution can change too—knowing which requests share a prefix lets you use different attention kernels that read the shared KV from HBM once instead of once per request.

Original post →

More from Infra

Infra channel →