Chinese Researchers Open-Source Cache-to-Cache: LLMs Talk Without Tokens, +14.2% Accuracy
Scobleizer · x · 2026-09-18
Chinese researchers open-sourced Cache-to-Cache (C2C), a paradigm letting LLMs communicate without generating a single word. Today, multi-agent systems must translate internal "thoughts" into text tokens, losing semantic richness and adding token-by-token latency. C2C instead uses a neural network to directly project and fuse the source model's KV-cache into the target model — pure semantic communication — with a learnable gating mechanism to pick which layers benefit most from cache transfer. Results: zero intermediate text-generation latency and accuracy gains of up to 14.2% over individual models.
More from Research
- Google publishes Dream RSI paper: agents must dream to recursively self-improve — Saboo_Shubham_ · 2026-09-18
- Researchers find S-Space: multimodal models encode left/right/up/down/depth in a linear subspace — jiqizhixin · 2026-09-18
- Economist proposes AI audits of pre-analysis plans to curb p-hacking in slow-moving fields — paulnovosad · 2026-09-18
- Google paper proposes Dream RSI: agents must dream to recursively self-improve — Saboo_Shubham_ · 2026-09-18
- Economists propose AI auditing of pre-analysis plan adherence to curb p-hacking — paulnovosad · 2026-09-18
- Stanford's 10-person Marin open lab is live-training a 535B model in the open — wandb · 2026-09-18