Tessera: retrieval-driven KV cache reuse cuts RAG serving TTFT by up to 3.6x

_reachsumit · x · 2026-09-29

A new arXiv paper, Tessera, presents an LLM serving system for RAG and retrieval-based agent memory. Motivation: the same retrieved content recurs across requests at different prompt positions and preceding contexts, defeating conventional prefix caching; the authors find that records recurring outside the matching prefix account for over 70% of injected memory tokens in agent-memory workloads. Composable KV reuse is possible, but online serving poses a management problem: a recurring unit's KV states may not exist yet, may be evicted, or may sit on another node.

Tessera makes retrieval the control plane for KV reuse: by exposing needed context units before model execution, it combines current demand with retrieval history, KV residency, and generation load to coordinate cache management and request routing. Generation nodes concurrently prepare locally cached, remotely cached, and missing states, retaining newly computed states off the critical path. Across RAG and agent-memory workloads, Tessera lowers mean TTFT by up to 3.6x over SGLang and LMCache with EPIC at matched request rates, sustains low TTFT where baselines saturate, and matches answer quality.

Original post →

More from Infra

Infra channel →