ThunderAgent (ICML 2026 Spotlight): Overcomes KV Cache Thrashing in Agentic Inference

togethercompute · x · 2026-07-30

Agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Running hundreds concurrently causes KV cache thrashing, where dumb LRU policies blindly evict caches needed seconds later, leading to cascading recomputations.

ThunderAgent fixes this at the scheduler level. It achieves 2.5x higher single-node throughput and 10x lower P50 latency at high concurrency. The paper was accepted as an ICML 2026 Spotlight.

Related event: ThunderAgent Doubles Agent Inference Throughput(6 posts)→

Original post →

More from coding & agent

coding & agent channel →