ReCache cuts agentic KV memory 92.43% and speeds first token 3.655x with minimal accuracy loss

techNmak · x · 2026-09-08

arXiv paper 2608.19662 (Yichu Fang et al.) proposes ReCache to fix prefix caching's failure on tool-augmented agents, where schemas recur in varying combinations:

On a benchmark from 7 public tool/skill datasets (including resource-disjoint tests), resource-wise attention matches dense invocation performance (82.3% vs 82.4% Inv-F1) with 3.655x time-to-first-token speedup; the full framework cuts allocated KV memory by 92.43% and speeds attention 1.423x. Code is open source.

Related event: ReCache Reuses Tool Schema KV Cache, Cutting VRAM by 92%(2 posts)→

Original post →

More from coding & agent

coding & agent channel →