CacheBack cuts 75% of agent-to-agent KV cache state, +14.7 points accuracy on FanOutQA

Maximillian Rossi · hf · 2026-10-05

Transferring full KV caches between agents avoids text decoding but memory grows linearly with context size and agent count. The authors propose receiver-conditioned communication: the receiver sends a short description of its information needs, which the sender uses to filter and compress its KV cache.

CacheBack is a training-free implementation based on the sender's attention weights. On FanOutQA with Qwen 3, it removes 75% of the state the receiving agent would otherwise get, improving accuracy by 14.7 percentage points and cutting median task-completion latency 3.2x versus text communication. Similar gains hold across dense Transformers, Mamba-attention hybrids, and sliding-window attention models.

Original post →

More from coding & agent

coding & agent channel →