ReCache Reuses Tool Schema KV Cache, Cutting VRAM by 92%

A new arXiv paper proposes ReCache, which builds independently reusable KV cache states per tool schema for LLM agents, saving 92% of VRAM and speeding up first-token generation by 3.7x.

2026-09-08 ~ 2026-09-08 · 2 related posts