Tsinghua's C2C Lets LLMs Skip Text and Merge KV-Caches Directly, 2.5x Faster with +14.2% Accuracy
anselm · x · 2026-09-19
Tsinghua and Infinigence teams open-sourced C2C (Cache-to-Cache), an ICLR 2026 paper that removes text from multi-agent LLM communication: a lightweight Neural Fuser rotates and grafts model A's KV-Cache directly into model B's, with a learnable gate deciding which layers absorb external cache. Results: 2.0–2.5x faster inference, up to 14.2% accuracy gain over a single model, and 5% over text-based agent collaboration — challenging the assumption that LLMs must talk in human language.
More from coding & agent
- Thorsten Ball demos agent picking its next command from shell history — IanArawjo · 2026-09-19
- LTX 2.5 native inpainting workflows with IC-Lora control released for free — No-Property3068 · 2026-09-19
- ZCode, Zhipu's coding app, silently uploads your entire Git history to Aliyun OSS — LocoMod · 2026-09-19
- DSPy crew ships lm15, a zero-dependency litellm alternative for multi-provider LLM calls — lateinteraction · 2026-09-19
- LangChain to host livestream on Jev, a model claiming 200x faster, 400x cheaper inference — LangChain · 2026-09-19
- Clairvoyance integrates Jev to help persistent-memory AI agents recall the right context — draginol · 2026-09-19