Tsinghua Open-Sources C2C: LLMs Communicate via KV-Cache, 2.5x Faster Inference

Tsinghua and Infinigence researchers open-sourced Cache-to-Cache (C2C), accepted by ICLR 2026, which lets LLM agents communicate directly via KV-Cache instead of text tokens, boosting inference speed by 2.5x.

2026-09-18 ~ 2026-09-19 · 2 related posts