CacheBack cuts 75% of agent-to-agent KV cache state, +14.7 points accuracy on FanOutQA
Maximillian Rossi · hf · 2026-10-05
Transferring full KV caches between agents avoids text decoding but memory grows linearly with context size and agent count. The authors propose receiver-conditioned communication: the receiver sends a short description of its information needs, which the sender uses to filter and compress its KV cache.
CacheBack is a training-free implementation based on the sender's attention weights. On FanOutQA with Qwen 3, it removes 75% of the state the receiving agent would otherwise get, improving accuracy by 14.7 percentage points and cutting median task-completion latency 3.2x versus text communication. Similar gains hold across dense Transformers, Mamba-attention hybrids, and sliding-window attention models.
More from coding & agent
- Context Language Models: letting AI edit its own memory with suffix cache reuse — jbhuang0604 · 2026-10-05
- tok_mcp: an MCP server that plans architecture docs to guide AI coding tools — modelcontextprotocol · 2026-10-05
- Hugging Face turns Claude Code, Codex and 10 coding harnesses into open RL environments — huggingface · 2026-10-05
- Music video FX made with Grok video, Wan, After Effects and Blender pipeline — Tranchillo · 2026-10-05
- Using four household devices for a local multi-agent setup: orchestrator plus workers — Jebbyk1 · 2026-10-05
- AI Gateway adds Exa, Ceramic and Linkup search providers with unified billing and observability — michellechen · 2026-10-05