Tessera: retrieval-driven KV cache reuse cuts RAG serving TTFT by up to 3.6x
_reachsumit · x · 2026-09-29
A new arXiv paper, Tessera, presents an LLM serving system for RAG and retrieval-based agent memory. Motivation: the same retrieved content recurs across requests at different prompt positions and preceding contexts, defeating conventional prefix caching; the authors find that records recurring outside the matching prefix account for over 70% of injected memory tokens in agent-memory workloads. Composable KV reuse is possible, but online serving poses a management problem: a recurring unit's KV states may not exist yet, may be evicted, or may sit on another node.
Tessera makes retrieval the control plane for KV reuse: by exposing needed context units before model execution, it combines current demand with retrieval history, KV residency, and generation load to coordinate cache management and request routing. Generation nodes concurrently prepare locally cached, remotely cached, and missing states, retaining newly computed states off the critical path. Across RAG and agent-memory workloads, Tessera lowers mean TTFT by up to 3.6x over SGLang and LMCache with EPIC at matched request rates, sustains low TTFT where baselines saturate, and matches answer quality.
More from Infra
- 23 of 128 Bittensor subnets now make real money; top GPU subnet billed $964K last month — bittingthembits · 2026-09-29
- Swapping matmul for associative-algebra layers boosts 110M LM throughput 7.8% — Ilya Koziev · 2026-09-29
- Shaw mocks data center opponents: hating compute while using the internet is incoherent — zealcaiden · 2026-09-29
- Modeling 1B agent VMs by 2030: what personal AI agents mean for CPU demand — AccBalanced · 2026-09-29
- Venice's tokenized inference-credit model is being copied — and oversupply looms — 0xJeff · 2026-09-29
- Model Casting: Mid-Training Recipe Sparsifies FFN Activations for Fewer FLOPs — francoisfleuret · 2026-09-29