Cache-to-Cache: LLMs talk directly via KV cache, skipping text

rochansinha · hn · 2026-09-19

An arXiv paper proposes Cache-to-Cache (C2C), letting LLMs communicate by mapping one model's KV cache into another's semantic space instead of generating text, reportedly cutting latency and improving multi-model collaboration over text-based communication.

Original post →

More from coding & agent

coding & agent channel →