Workload-aware inference: why batch LLM pipelines should plan queries like databases do
sh_reya · x · 2026-10-02
The author is bullish on workload-aware inference, arguing database systems are the crown jewel of exploiting workload knowledge for performance—and any system with a declarative interface and budding LLM inference demand is a good candidate.
In a batch data processing pipeline with multiple LLM operations, the typical API approach makes one call per row per operation, served independently. If you know all requests up front, you can plan the whole job: reorder requests within an operation to share KV cache, run the most selective filter first, and reuse each document's KV cache across operations. Beyond planning, execution can change too—knowing which requests share a prefix lets you use different attention kernels that read the shared KV from HBM once instead of once per request.
More from Infra
- Fireworks shows numerical mismatch can collapse RL training in 25 steps on GLM and MoE models — sophiamyang · 2026-10-02
- $13B Baseten bets on open models, launches Base Labs research lab — baseten · 2026-10-02
- Kaigen details Unity benchmark setup: Runtime Speed with native C/C++ multithreading — gdechichi · 2026-10-02
- a16z: every $100 into AI buildout sends $50 to chips, $20 to power; 100+ charts — demian_ai · 2026-10-02
- Where does GPU spend actually go: training, inference, or idle reserved capacity? — Borges_Engineer · 2026-10-02
- HBM4 controllers eat ~16% of Nvidia's Rubin die, sparking optical interconnect debate — BenBajarin · 2026-10-02