ThunderAgent (ICML 2026 Spotlight): Overcomes KV Cache Thrashing in Agentic Inference
togethercompute · x · 2026-07-30
Agent workflows alternate between GPU-heavy reasoning and GPU-idle waiting on tools. Running hundreds concurrently causes KV cache thrashing, where dumb LRU policies blindly evict caches needed seconds later, leading to cascading recomputations.
ThunderAgent fixes this at the scheduler level. It achieves 2.5x higher single-node throughput and 10x lower P50 latency at high concurrency. The paper was accepted as an ICML 2026 Spotlight.
Related event: ThunderAgent Doubles Agent Inference Throughput(6 posts)→
More from coding & agent
- AI Bug Fixes Often Incomplete, Creating Messy Open Source Security — curious_vii · 2026-07-30
- Codex Tip: Turn Prompts into Repeatable, Scriptable Workflows — reach_vb · 2026-07-30
- ShadKit Open Source: Porting shadcn/ui and Vercel AI Elements to SwiftUI — jasonkneen · 2026-07-30
- Test: Enabling Two API Settings Boosts GPT-5.6 ARC-AGI-3 Score by 3x — sandersted · 2026-07-30
- dspy-monty-interpreter v0.3.0 Released: Multithreading and Isolated Execution — dbreunig · 2026-07-30
- Current AI Agent Memory Systems Are Just Hacky RAG Wrappers — Trick_Stretch_4746 · 2026-07-30