ReCache reuses tool-schema KV states, cutting agent memory 92% with little accuracy loss

techNmak · x · 2026-09-08

ReCache is a KV-cache reuse and compression method for tool-augmented LLM agents: it builds independently reusable KV states per tool/skill schema with resource-local positions, then applies resource-wise attention, structural routing over layer/KV-head groups, and semantic pruning that keeps only invocation-critical fields. See the companion post with the arXiv link for full numbers.

Related event: ReCache Reuses Tool Schema KV Cache, Cutting VRAM by 92%(2 posts)→

Original post →

More from coding & agent

coding & agent channel →