Infra optimization potential when controlling open model inference

_ScottCondron · x · 2026-09-02

The post argues that while most model harnesses are similar to preserve KV cache hit rates for proprietary models, controlling the underlying infrastructure with open models allows for deeper optimization. Potential optimizations include session affinity, speculative tool calling, workload-based request priorities, prefetching KV cache during tool calls, colocating subagents with tool calls, and prewarming tool calling sandboxes.

Original post →

More from Infra

Infra channel →